ML-QuantSubscribe

Machine learningML & AI Methods

Critique-out-Loud Reward Models

The article presents CLoud reward models that use human feedback to improve reinforcement learning, enhancing accuracy and win rate in ArenaHard.

Featured in No. 63 on 28 Aug 2024 · 7 days after release · 99 citations today

Released
21 Aug 2024
First featured
No. 63 · 28 Aug 2024
Citations (Semantic Scholar)
99
Influential citations
6
Published in
Not yet, as far as Semantic Scholar knows
Shares when featured
46
Identifier
arXiv:2408.11791

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page