Machine learningLLMs & Text
AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration
The study suggests Activation-aware Weight Quantization (AWQ), a hardware-friendly method for quantizing large language models that reduces error and improves performance on various benchmarks.
Featured in No. 58 on 24 Jul 2024 · · 1,800 citations today · published in GetMobile: Mobile Computing and Communications
- Released
- 1 Jun 2023
- First featured
- No. 58 · 24 Jul 2024
- Citations (Semantic Scholar)
- 1,800
- Influential citations
- 268
- Published in
- GetMobile: Mobile Computing and Communications
- Shares when featured
- 146
- Identifier
- arXiv:2306.00978
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).