Machine learningML & AI Methods
DPO Meets PPO: Reinforced Token Optimization for RLHF
A new framework is introduced that models Reinforcement Learning from Human Feedback as a Markov decision process, using an algorithm that learns from preference data.
Featured in No. 47 on 1 May 2024 · 2 days after release · 143 citations today · published in International Conference on Machine Learning
- Released
- 29 Apr 2024
- First featured
- No. 47 · 1 May 2024
- Citations (Semantic Scholar)
- 143
- Influential citations
- 11
- Published in
- International Conference on Machine Learning
- Shares when featured
- 9
- Identifier
- arXiv:2404.18922
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).