Machine learningLLMs & Text
Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference
The piece presents Dynamic Memory Compression (DMC), a method for compressing key-value cache in large language models that increases throughput and maintains performance while accommodating larger contexts and batches within a given memory budget.
Featured in No. 58 on 24 Jul 2024 · · 133 citations today · published in International Conference on Machine Learning
- Released
- 14 Mar 2024
- First featured
- No. 58 · 24 Jul 2024
- Citations (Semantic Scholar)
- 133
- Influential citations
- 2
- Published in
- International Conference on Machine Learning
- Shares when featured
- 106
- Identifier
- arXiv:2403.09636
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).