ML-QuantSubscribe

Machine learningLLMs & Text

KTO: Model Alignment as Prospect Theoretic Optimization

The study presents KTO, a new method that uses a Kahneman-Tversky model to align language models with human feedback, performing better than preference-based methods by learning from a binary signal of output desirability.

Featured in No. 76 on 27 Nov 2024 · · 162 citations today

Released
2 Feb 2024
First featured
No. 76 · 27 Nov 2024
Citations (Semantic Scholar)
162
Influential citations
22
Published in
Not yet, as far as Semantic Scholar knows
Shares when featured
71
Identifier
arXiv:2402.01306

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page