Machine learningML & AI Methods
Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
The article introduces Cross-Layer Attention (CLA), a new attention design that minimizes the key-value cache size, allowing for longer sequence lengths and larger batch sizes during inference.
Featured in No. 51 on 28 May 2024 · 7 days after release · 140 citations today · published in Neural Information Processing Systems
- Released
- 21 May 2024
- First featured
- No. 51 · 28 May 2024
- Citations (Semantic Scholar)
- 140
- Influential citations
- 10
- Published in
- Neural Information Processing Systems
- Shares when featured
- 206
- Identifier
- arXiv:2405.12981
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).