ML-QuantSubscribe

Machine learningLLMs & Text

Squeezed Attention: Accelerating Long Context Length LLM Inference

LLM Inference: Squeezed Attention is a proposed method to speed up Large Language Model applications by using K-means clustering to group similar keys in fixed context inputs, reducing computational costs and enhancing inference efficiency.

Featured in No. 75 on 20 Nov 2024 · 6 days after release · 60 citations today · published in Annual Meeting of the Association for Computational Linguistics

Released
14 Nov 2024
First featured
No. 75 · 20 Nov 2024
Citations (Semantic Scholar)
60
Influential citations
3
Published in
Annual Meeting of the Association for Computational Linguistics
Shares when featured
25
Identifier
arXiv:2411.09688

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page