Machine learningLLMs & Text
Squeezed Attention: Accelerating Long Context Length LLM Inference
LLM Inference: Squeezed Attention is a proposed method to speed up Large Language Model applications by using K-means clustering to group similar keys in fixed context inputs, reducing computational costs and enhancing inference efficiency.
Featured in No. 75 on 20 Nov 2024 · 6 days after release · 60 citations today · published in Annual Meeting of the Association for Computational Linguistics
- Released
- 14 Nov 2024
- First featured
- No. 75 · 20 Nov 2024
- Citations (Semantic Scholar)
- 60
- Influential citations
- 3
- Published in
- Annual Meeting of the Association for Computational Linguistics
- Shares when featured
- 25
- Identifier
- arXiv:2411.09688
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).