Machine learningLLMs & Text
Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF
The Exploratory Preference Optimization (XPO) algorithm has been introduced for online exploration in Reinforcement Learning from Human Feedback (RLHF), potentially enhancing language model training.
Featured in No. 52 on 5 Jun 2024 · 5 days after release · 115 citations today
- Released
- 31 May 2024
- First featured
- No. 52 · 5 Jun 2024
- Citations (Semantic Scholar)
- 115
- Influential citations
- 11
- Published in
- Not yet, as far as Semantic Scholar knows
- Shares when featured
- 7
- Identifier
- arXiv:2405.21046
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).