---
title: LLM Pruning and Distillation in Practice: The Minitron Approach
url: https://www.ml-quant.com/papers/arxiv/2408.11796/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2408.11796
source_url: https://arxiv.org/abs/2408.11796
featured: 2024-08-28
citations: 111
topic: LLMs & Text
---


# LLM Pruning and Distillation in Practice: The Minitron Approach

The article discusses the successful compression of Llama 3.1 8B and MistralNeMo 12B models to smaller parameters using pruning and distillation strategies, with the results tested on common benchmarks and the base model weights made available on Hugging Face.

- Source: https://arxiv.org/abs/2408.11796
- Identifier: arXiv:2408.11796
- Released: 2024-08-21
- First featured: Quant Letter No. 63 (2024-08-28): https://www.ml-quant.com/issues/2024-08-28/
- Citations (Semantic Scholar): 111
- Published in: not yet
- Topic: LLMs & Text

## Related

- [Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters](https://www.ml-quant.com/papers/arxiv/2408.03314/): The research investigates enhancing Large Language Models' (LLMs) performance using more test-time computation, suggesting a compute-optimal scaling strategy based on prompt difficulty.
- [AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration](https://www.ml-quant.com/papers/arxiv/2306.00978/): The study suggests Activation-aware Weight Quantization (AWQ), a hardware-friendly method for quantizing large language models that reduces error and improves performance on various benchmarks.
- [Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling](https://www.ml-quant.com/papers/arxiv/2412.05271/): The paper presents InternVL 2.5, a sophisticated multimodal large language model that performs well on various benchmarks, exceeding 70% on the MMMU benchmark.
- [MemGPT: Towards LLMs as Operating Systems](https://www.ml-quant.com/papers/arxiv/2310.08560/): Extended Context in LLMs: MemGPT is a system that manages different memory levels, providing extended context within large language models' limited context windows, enhancing document analysis and multi-session chat performance.
- [SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models](https://www.ml-quant.com/papers/arxiv/2303.08896/): Hallucination Detection for LLMs: The paper presents SelfCheckGPT, a new approach for fact-checking black-box model responses without an external database, proving its superior ability to detect and rank factual and non-factual sentences.
- [A Simple and Effective Pruning Approach for Large Language Models](https://www.ml-quant.com/papers/arxiv/2306.11695/): Wanda, a new method, efficiently prunes weights in Large Language Models without retraining, offering a more efficient approach to inducing sparsity in pretrained models.
