KV Cache Engineering for LLM Serving
KV cache engineering is critical for managing GPU memory during LLM inference because the cache grows dynamically with sequence length. The article outlines twelve techniques categorized by whether they reduce heads, layers, tokens, dimensions, or bits. Architectural methods like GQA and MLA modify the model structure, while serving-level optimizations like paging and Quest improve memory handling and read efficiency. These techniques involve trade-offs between memory savings, computational speed, and model accuracy.
閱讀原文 ↗KV cache engineering for LLM serving
KV cache engineering is critical for managing GPU memory during LLM inference because the cache grows dynamically with sequence length. The article outlines twelve techniques categorized by whether they reduce heads, layers, tokens, dimensions, or bits. Architectural methods like GQA and MLA modify the model structure, while serving-level optimizations like paging and Quest improve memory handling and read efficiency. These techniques involve trade-offs between memory savings, computational speed, and model accuracy.
- Llama 3.1 70B requires approximately 40 GB of BF16 KV cache for a single 128K-token sequence.
- GQA and MQA reduce memory usage by sharing key-value heads across multiple query heads within a layer.
- Cross-layer attention (CLA) can achieve a 2x reduction in KV cache by sharing cached data between adjacent transformer layers.
- Sliding-window attention fixes the cache size for local layers by only retaining the most recent tokens.
- Multi-head latent attention (MLA), introduced by DeepSeek-V2, compresses the KV cache into a latent representation to increase throughput.
- Hybrid models like Jamba and Qwen3-Next interleave attention layers with recurrent layers to maintain fixed-size states.
- Quest optimizes inference speed by reading only the most relevant pages of the cache based on query scores without deleting entries.