ML-QuantSubscribe

Machine learningLLMs & Text

Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference

The piece presents Dynamic Memory Compression (DMC), a method for compressing key-value cache in large language models that increases throughput and maintains performance while accommodating larger contexts and batches within a given memory budget.

Featured in No. 58 on 24 Jul 2024 · · 133 citations today · published in International Conference on Machine Learning

Released
14 Mar 2024
First featured
No. 58 · 24 Jul 2024
Citations (Semantic Scholar)
133
Influential citations
2
Published in
International Conference on Machine Learning
Shares when featured
106
Identifier
arXiv:2403.09636

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page