Machine learningLLMs & Text
LLM Pruning and Distillation in Practice: The Minitron Approach
The article discusses the successful compression of Llama 3.1 8B and MistralNeMo 12B models to smaller parameters using pruning and distillation strategies, with the results tested on common benchmarks and the base model weights made available on Hugging Face.
Featured in No. 63 on 28 Aug 2024 · 7 days after release · 111 citations today
- Released
- 21 Aug 2024
- First featured
- No. 63 · 28 Aug 2024
- Citations (Semantic Scholar)
- 111
- Influential citations
- 11
- Published in
- Not yet, as far as Semantic Scholar knows
- Shares when featured
- 372
- Identifier
- arXiv:2408.11796
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).