Why An LLM’s Memory Gets Expensive and How to Fix It
To generate each token, a model executes an attention step comparing the newest token against earlier tokens using key and value vectors. Recomputing these vectors for every preceding token at each step is computationally wasteful because their values remain static. The KV cache resolves this inefficiency by storing computed key and value vectors so that only the new token requires computation. However, reading this cache on every step introduces a new cost bottleneck and can cause unexpected out-of-memory errors because the cache retains high-dimensional vectors rather than text. Token generation operates in two distinct phases with different hardware bottlenecks: prefill and decoding. During prefill, all input tokens are processed in parallel to construct the key-value cache, making this phase compute-bound and limited by arithmetic throughput. Decoding generates tokens autoregressively one by one, requiring a full read of the key-value cache from GPU memory for every step, which makes it memory-bound. As a result, the primary bottleneck in long-context generation stems from repeatedly streaming the entire cache across the memory bus rather than simply storing it. The key-value cache size of a large language model is determined by multiplying the number of layers, key-value heads, head dimension, bytes per number, context length, batch size, and a factor of two. Cache requirements scale linearly with both token count and batch size, making long contexts and concurrent users resource-intensive. For instance, a single 128,000-token request on Llama 3 70B demands approximately 40 gigabytes of memory, consuming half of an 80-gigabyte GPU card. Optimization techniques aim to mitigate memory demands by reducing individual factors in this multiplication.
閱讀原文 ↗目錄
Recomputation
To generate each token, a model executes an attention step comparing the newest token against earlier tokens using key and value vectors. Recomputing these vectors for every preceding token at each step is computationally wasteful because their values remain static. The KV cache resolves this inefficiency by storing computed key and value vectors so that only the new token requires computation. However, reading this cache on every step introduces a new cost bottleneck and can cause unexpected out-of-memory errors because the cache retains high-dimensional vectors rather than text.
- Generating a token requires an attention step where the newest token is compared against prior tokens using key and value vectors.
- Recomputing prior keys and values at each step causes work per token to increase unnecessarily as context grows.
- The KV cache eliminates redundant computation by storing key and value vectors after their initial calculation.
- Reading the KV cache on every step constitutes an increasing operational cost.
- Because the KV cache stores vector representations rather than raw text, it can cause out-of-memory errors even when the model weights fit comfortably in memory.
Decoding
Token generation operates in two distinct phases with different hardware bottlenecks: prefill and decoding. During prefill, all input tokens are processed in parallel to construct the key-value cache, making this phase compute-bound and limited by arithmetic throughput. Decoding generates tokens autoregressively one by one, requiring a full read of the key-value cache from GPU memory for every step, which makes it memory-bound. As a result, the primary bottleneck in long-context generation stems from repeatedly streaming the entire cache across the memory bus rather than simply storing it.
- Token generation is split into two phases: prefill and decoding.
- The prefill phase processes all input tokens in parallel and is compute-bound, limited by GPU arithmetic capacity.
- The decoding phase produces one token at a time and is memory-bound, limited by memory bandwidth.
- Each decoding step requires reading all stored key and value vectors out of GPU memory.
- Long-context performance degradation is driven by the cost of sweeping through the cache on every generated token rather than memory capacity alone.
Scaling
The key-value cache size of a large language model is determined by multiplying the number of layers, key-value heads, head dimension, bytes per number, context length, batch size, and a factor of two. Cache requirements scale linearly with both token count and batch size, making long contexts and concurrent users resource-intensive. For instance, a single 128,000-token request on Llama 3 70B demands approximately 40 gigabytes of memory, consuming half of an 80-gigabyte GPU card. Optimization techniques aim to mitigate memory demands by reducing individual factors in this multiplication.
- The KV cache size equals 2 times layers times key-value heads times head dimension times bytes per number times tokens times batch size.
- Cache memory usage grows linearly with token count, meaning doubling context length doubles cache consumption.
- Cache memory scales linearly with batch size, proportionally increasing with concurrent requests.
- A single 128,000-token request on Llama 3 70B requires approximately 40 gigabytes of cache memory.
- Llama 3 70B architecture parameters include 80 layers, 8 key-value heads, a head dimension of 128, and 2 bytes per number.
Attention
Attention architecture modifications can substantially reduce the key-value cache memory footprint per token, though these designs must be committed to during training. Grouped-query attention saves memory by having multiple query heads share key-value heads, offering a stable middle ground compared to aggressive multi-query attention. Multi-head latent attention further compresses cache demands by projecting key and value representations into a smaller latent space before caching, as demonstrated by DeepSeek-V3. While these architectural techniques reduce memory traffic at large scale, they require trade-offs in serving overhead, quality, or implementation complexity.
- Grouped-query attention lets several query heads share one key-value head, reducing cache memory roughly eightfold in models like Llama 2 70B, Llama 3 70B, and Mistral 7B.
- Multi-query attention achieves maximum memory savings by sharing a single key-value head across all query heads, but it often causes training instability and quality degradation.
- Multi-head latent attention, introduced in DeepSeek models, compresses keys and values into a smaller latent representation before caching and decompresses them on read.
- DeepSeek-V3 achieves a cache footprint of approximately 70 kilobytes per token via latent attention, compared to 192 to 328 kilobytes per token in comparable grouped-query attention models.
- Head-sharing and latent attention techniques require architectural control during training rather than being applicable post-training.
Quantization
Quantization reduces the memory footprint of the key-value cache by converting numbers from standard 16-bit precision down to 8-bit or 4-bit representations. This technique requires no model retraining, cutting cache size in half at 8 bits and in half again at 4 bits. While 8-bit quantization typically produces negligible accuracy drops under one percent, 4-bit storage shows measurable degradation on complex tasks like multi-needle retrieval. Although specialized algorithms outperform plain rounding by preserving precision for key values, simple rounding still captures the majority of compression gains.
- Quantization decreases memory usage by reducing key and value storage precision from 16 bits to 8 bits or 4 bits.
- Transitioning from 16-bit to 8-bit precision halves the cache size, and moving to 4-bit halves it again.
- Quantization can be applied directly to existing models without retraining.
- Eight-bit storage typically loses well under one percent of accuracy for most standard workloads.
- Four-bit storage introduces measurable performance degradation in demanding tasks such as multi-needle retrieval.
- Specialized quantization methods outperform plain rounding by maintaining higher precision for outlier numbers in the cache.
Eviction
Eviction reduces token counts in language models by discarding entries deemed unlikely to be needed. A typical implementation maintains both the most recent context window and a few opening tokens, which function as attention anchors to stabilize model generation. However, eviction faces a structural flaw because the model cannot predict whether a discarded token will be essential for subsequent queries. As a result, aggressive eviction techniques often fail on retrieval tasks where key information is buried mid-sequence.
- Eviction lowers memory usage by discarding tokens the model is predicted not to need.
- A common eviction strategy retains recent tokens alongside a small set of opening sequence tokens.
- Opening tokens absorb significant attention regardless of content, serving as stability anchors for generation.
- The fundamental challenge of eviction is that token importance depends on future, unknown inputs.
- Aggressive cache eviction disproportionately harms retrieval tasks by dropping details located in the middle of documents.
- Importance-scoring schemes attempt to predict safe drops, but they do not eliminate the structural limitation.
Serving
Serving systems historically suffered high memory fragmentation by reserving contiguous blocks sized for worst-case output lengths. Paged attention addresses this by splitting the cache into small fixed-size pages allocated on demand, reducing fragmentation from 60-80% to under 4% and boosting throughput by two to three times. Furthermore, storing cache in shareable blocks enables prefix caching, allowing requests with identical prefixes to point to the same physical memory. Providers such as OpenAI and Anthropic report 50 to 90 percent cost and latency reductions on cache hits, though cross-user cache sharing introduces timing side-channel security risks.
- Older serving architectures wasted 60 to 80 percent of cache memory on fragmentation by allocating contiguous blocks sized for maximum potential output length.
- Paged attention adapts operating system memory paging to LLM caches, dropping fragmentation below 4 percent and doubling or tripling throughput.
- Shareable cache blocks serve as the foundation for prefix caching (productized by major APIs as prompt caching).
- Prefix caching yields 50 to 90 percent reductions in cost and latency for repetitive workloads like multi-thousand-token system prompts according to OpenAI and Anthropic.
- Sharing cached memory blocks across users introduces potential timing side-channel vulnerabilities that can leak prompt data.
Tradeoffs
Different memory-saving techniques for model caches offer varied tradeoffs between memory reduction, implementation complexity, and output quality. Methods like grouped-query attention, paged attention, and prefix caching provide memory benefits with virtually no quality loss. In contrast, techniques such as low-bit quantization, latent attention, and eviction introduce notable quality risks or require significant engineering effort. Ultimately, these optimizations become vital in long-context and high-concurrency settings where cache overhead dominates, requiring strategies to be matched to specific workload patterns.
- Grouped-query attention serves as a near-costless default that ships with almost every modern model.
- Paged attention and prefix caching preserve output quality by altering how the cache is stored and shared rather than changing its content.
- Quantization remains safe at 8 bits but increases the risk of quality degradation at 4 bits and below.
- Latent attention yields substantial architectural savings but demands notable engineering effort to serve effectively.
- Eviction frees large amounts of memory but risks discarding information required for subsequent token generation.
- Cache reduction techniques are primarily necessary in long-context and high-concurrency workloads where cache memory becomes the dominant bottleneck.
Conclusion
The cost of long-context inference is fundamentally determined by the size of the attention cache. Because decoding reads the entire cache for every generated token, the cache creates both storage and memory-bandwidth bottlenecks. Shrinking the cache improves generation speed, and optimization techniques target different aspects of cache size and management. Approaches include reducing per-token cost, quantizing values, evicting tokens, and improving memory layout and sharing.
- Long-context inference cost is dictated by cache size, which imposes both storage and memory bandwidth overheads.
- During decoding, the full cache must be read on every token generation step.
- Grouped query attention and latent attention decrease the memory cost associated with each token.
- Quantization reduces memory requirements by encoding values in fewer bits.
- Cache eviction reduces the total number of retained tokens.
- Paged attention and prefix caching enhance cache management and enable sharing across requests.