Machine learningLLMs & Text
You Only Cache Once: Decoder-Decoder Architectures for Language Models
YOCO architecture improves large language models by reducing GPU memory usage and speeding up the prefill stage, outperforming the Transformer model.
Featured in No. 49 on 15 May 2024 · 7 days after release · 155 citations today · published in Neural Information Processing Systems
- Released
- 8 May 2024
- First featured
- No. 49 · 15 May 2024
- Citations (Semantic Scholar)
- 155
- Influential citations
- 14
- Published in
- Neural Information Processing Systems
- Shares when featured
- 272
- Identifier
- arXiv:2405.05254
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).