ML-QuantSubscribe

Machine learningLLMs & Text

Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF

The Exploratory Preference Optimization (XPO) algorithm has been introduced for online exploration in Reinforcement Learning from Human Feedback (RLHF), potentially enhancing language model training.

Featured in No. 52 on 5 Jun 2024 · 5 days after release · 115 citations today

Released
31 May 2024
First featured
No. 52 · 5 Jun 2024
Citations (Semantic Scholar)
115
Influential citations
11
Published in
Not yet, as far as Semantic Scholar knows
Shares when featured
7
Identifier
arXiv:2405.21046

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page