ML-QuantSubscribe

Machine learningML & AI Methods

SimPO: Simple Preference Optimization with a Reference-Free Reward

Simple Preference Optimization: SimPO improves reinforcement learning from human feedback by using the average log probability of a sequence as the implicit reward, enhancing training stability and computational efficiency.

Featured in No. 56 on 10 Jul 2024 · 48 days after release · 1,173 citations today · published in Neural Information Processing Systems

Released
23 May 2024
First featured
No. 56 · 10 Jul 2024
Citations (Semantic Scholar)
1,173
Influential citations
232
Published in
Neural Information Processing Systems
Shares when featured
728
Identifier
arXiv:2405.14734

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page