Machine learningML & AI Methods
SimPO: Simple Preference Optimization with a Reference-Free Reward
Simple Preference Optimization: SimPO improves reinforcement learning from human feedback by using the average log probability of a sequence as the implicit reward, enhancing training stability and computational efficiency.
Featured in No. 56 on 10 Jul 2024 · 48 days after release · 1,173 citations today · published in Neural Information Processing Systems
- Released
- 23 May 2024
- First featured
- No. 56 · 10 Jul 2024
- Citations (Semantic Scholar)
- 1,173
- Influential citations
- 232
- Published in
- Neural Information Processing Systems
- Shares when featured
- 728
- Identifier
- arXiv:2405.14734
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).