ML-QuantSubscribe

Machine learningLLMs & Text

LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation

SimLayerKV is a technique that minimizes memory usage in large language models by identifying and reducing cache in lazy layers, achieving significant cache compression with minimal performance loss.

Featured in No. 71 on 23 Oct 2024 · 6 days after release · 13 citations today · published in Trans. Mach. Learn. Res.

Released
17 Oct 2024
First featured
No. 71 · 23 Oct 2024
Citations (Semantic Scholar)
13
Influential citations
0
Published in
Trans. Mach. Learn. Res.
Shares when featured
8
Identifier
arXiv:2410.13846

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page