Machine learningML & AI Methods
Dataset Reset Policy Optimization for RLHF
The DR-PO algorithm enhances Reinforcement Learning by incorporating offline preference data into online policy training, outperforming other techniques in summarization and the Anthropic Helpful Harmful dataset.
Featured in No. 45 on 17 Apr 2024 · 5 days after release · 44 citations today
- Released
- 12 Apr 2024
- First featured
- No. 45 · 17 Apr 2024
- Citations (Semantic Scholar)
- 44
- Influential citations
- 7
- Published in
- Not yet, as far as Semantic Scholar knows
- Shares when featured
- 23
- Identifier
- arXiv:2404.08495
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).