ML-QuantSubscribe

Machine learningLLMs & Text

RLPF: Reinforcement Learning from Prediction Feedback for User Summarization with LLMs

Reinforcement Learning from Prediction Feedback (RLPF) is a method that refines Large Language Models to produce concise, human-readable summaries, enhancing task performance and summary quality.

Featured in No. 65 on 10 Sep 2024 · 4 days after release · 16 citations today · published in AAAI Conference on Artificial Intelligence

Released
6 Sep 2024
First featured
No. 65 · 10 Sep 2024
Citations (Semantic Scholar)
16
Influential citations
3
Published in
AAAI Conference on Artificial Intelligence
Shares when featured
9
Identifier
arXiv:2409.04421

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page