Machine learningML & AI Methods
Critique-out-Loud Reward Models
The article presents CLoud reward models that use human feedback to improve reinforcement learning, enhancing accuracy and win rate in ArenaHard.
Featured in No. 63 on 28 Aug 2024 · 7 days after release · 99 citations today
- Released
- 21 Aug 2024
- First featured
- No. 63 · 28 Aug 2024
- Citations (Semantic Scholar)
- 99
- Influential citations
- 6
- Published in
- Not yet, as far as Semantic Scholar knows
- Shares when featured
- 46
- Identifier
- arXiv:2408.11791
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).