Machine learningLLMs & Text
SALMON: Self-Alignment with Instructable Reward Models
Minimal Human Supervision Language Model Alignment: The paper introduces SALMON, a new method for aligning base language models with minimal human supervision using principle-following reward models, showing its superior performance on multiple benchmark datasets.
Featured in No. 21 on 16 Oct 2023 · 7 days after release · 67 citations today · published in International Conference on Learning Representations
- Released
- 9 Oct 2023
- First featured
- No. 21 · 16 Oct 2023
- Citations (Semantic Scholar)
- 67
- Influential citations
- 7
- Published in
- International Conference on Learning Representations
- Shares when featured
- 36
- Identifier
- arXiv:2310.05910
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).