ML-QuantSubscribe

Machine learningLLMs & Text

From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function

The research explores Direct Preference Optimization (DPO) in Reinforcement Learning From Human Feedback (RLHF), showing its ability to assign credit and its similarity to search-based algorithms in language generation.

Featured in No. 46 on 24 Apr 2024 · 6 days after release · 273 citations today

Released
18 Apr 2024
First featured
No. 46 · 24 Apr 2024
Citations (Semantic Scholar)
273
Influential citations
27
Published in
Not yet, as far as Semantic Scholar knows
Shares when featured
107
Identifier
arXiv:2404.12358

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page