arXivML & AI Methods
VickreyFeedback: Cost-efficient Data Construction for Reinforcement Learning from Human Feedback
An auction mechanism is introduced to enhance cost-efficiency in fine-tuning large language models using Reinforcement Learning from Human Feedback, focusing on quality feedback and model performance.
Featured in No. 79 on 18 Dec 2024 · · 3 citations today · published in Prima
- Released
- 27 Sep 2024
- First featured
- No. 79 · 18 Dec 2024
- Citations (Semantic Scholar)
- 3
- Influential citations
- 1
- Published in
- Prima
- Shares when featured
- 14
- Identifier
- arXiv:2409.18417
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).