Machine learningLLMs & Text
Self-Play Preference Optimization for Language Model Alignment
The article suggests a self-play-based method, SPPO, for language model alignment, which can effectively enhance the likelihood of the selected response and reduce that of the discarded one.
Featured in No. 51 on 28 May 2024 · 27 days after release · 278 citations today · published in International Conference on Learning Representations
- Released
- 1 May 2024
- First featured
- No. 51 · 28 May 2024
- Citations (Semantic Scholar)
- 278
- Influential citations
- 41
- Published in
- International Conference on Learning Representations
- Shares when featured
- 309
- Identifier
- arXiv:2405.00675
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).