How to Make LLMs 3X Faster
Text generation in language models operates autoregressively, producing content one token at a time through sequential forward passes across all model layers. Because each new token strictly depends on the presence of all previous tokens, forward passes cannot be executed simultaneously without breaking coherence. Consequently, total generation duration equals the token count multiplied by the per-pass latency, which scales with model size. Modern inference systems mitigate the computational burden of subsequent passes by using a KV cache to store prior attention states, though the requirement for one pass per token remains. During model inference, a forward pass spends the majority of its execution time moving data rather than computing arithmetic. In token generation, model weights must be transferred from VRAM to compute units for each single token, resulting in low compute utilization between 20 and 40 percent while the memory bus runs near capacity. In contrast, prompt processing achieves 90 to 95 percent utilization because weights are reused across thousands of input tokens. Consequently, increasing memory bandwidth accelerates token generation significantly more than adding raw compute capacity. Transformers can evaluate an entire sequence simultaneously in a single forward pass, generating a next-token prediction for every position at once. Causal masking within the attention mechanism ensures each position only attends to preceding tokens, maintaining the validity of autoregressive conditioning. Consequently, candidate tokens can be verified in parallel during a single pass rather than sequentially, mirroring the computational efficiency of prompt processing.
閱讀原文 ↗目錄
Autoregressive Decoding
Text generation in language models operates autoregressively, producing content one token at a time through sequential forward passes across all model layers. Because each new token strictly depends on the presence of all previous tokens, forward passes cannot be executed simultaneously without breaking coherence. Consequently, total generation duration equals the token count multiplied by the per-pass latency, which scales with model size. Modern inference systems mitigate the computational burden of subsequent passes by using a KV cache to store prior attention states, though the requirement for one pass per token remains.
- Text generation runs sequentially, computing a probability distribution and appending one token at a time.
- A 500-token output necessitates 500 sequential forward passes because each token depends on the previous ones.
- Total generation latency is determined by multiplying the number of generated tokens by the duration of a single forward pass.
- Per-token cost is uniform regardless of whether the output is text or code, but increases on larger models running on identical hardware.
- A KV cache stores attention states for previously processed tokens, reducing per-pass workload without eliminating the pass-per-token requirement.
Memory Bandwidth
During model inference, a forward pass spends the majority of its execution time moving data rather than computing arithmetic. In token generation, model weights must be transferred from VRAM to compute units for each single token, resulting in low compute utilization between 20 and 40 percent while the memory bus runs near capacity. In contrast, prompt processing achieves 90 to 95 percent utilization because weights are reused across thousands of input tokens. Consequently, increasing memory bandwidth accelerates token generation significantly more than adding raw compute capacity.
- A single forward pass spends most of its duration moving data from memory rather than executing arithmetic.
- For a 70-billion-parameter model at 16-bit precision, approximately 140 GB of weights must be transferred per generated token.
- Compute utilization falls from 90-95% during prompt processing to 20-40% during token generation.
- Prompt processing applies loaded weights across thousands of tokens at once, while token generation reloads weights for a single token.
- GPU memory bandwidth is a more significant driver of token generation speed than raw compute capability.
Parallel Verification
Transformers can evaluate an entire sequence simultaneously in a single forward pass, generating a next-token prediction for every position at once. Causal masking within the attention mechanism ensures each position only attends to preceding tokens, maintaining the validity of autoregressive conditioning. Consequently, candidate tokens can be verified in parallel during a single pass rather than sequentially, mirroring the computational efficiency of prompt processing.
- A Transformer evaluates all sequence positions simultaneously within a single forward pass, generating a next-token prediction at each position.
- Causal masking inside the attention mechanism prevents positions from attending to subsequent tokens, preserving identical conditioning to sequential generation.
- Prompt processing exploits this parallel forward pass to process thousands of tokens in one pass instead of one token at a time.
- Candidate tokens can be verified simultaneously by appending them to the context and running a single forward pass.
- Verification and generation perform identical computational work per position, but parallel verification saves cost by grouping evaluations into one pass.
Draft and Verify
The draft and verify framework accelerates language model inference by combining a small draft model with a large target model. In each round, the draft model serially generates a sequence of candidate tokens, which the target model verifies simultaneously in a single forward pass. Tokens are accepted sequentially until the first mismatch, at which point the target model's already-computed prediction is retained as an additional correct token. This setup bounds the worst-case throughput to that of standard single-token decoding while offering significant speedups with draft lengths typically set between 3 and 5.
- The inference loop relies on two models: a primary target model and a draft model that has 10 to 20 times fewer parameters and typically shares the same tokenizer.
- Each round executes in three steps: serial drafting of K tokens, single-pass batched evaluation by the target model, and left-to-right verification.
- Matching tokens are retained, and the target model's prediction at the first mismatch position is accepted at no extra computation cost.
- In the worst case where all candidate tokens are rejected, the system still outputs one correct token, matching the baseline output of standard decoding.
- Draft length K is commonly tuned between 3 and 5 to balance potential speedups against the diminishing probability of later tokens being accepted.
Lossless Guarantee
Speculative decoding generates text with the exact same statistical properties as the target model operating independently, enforced via an acceptance rule. In greedy decoding, candidate tokens are accepted only if they match the target model's top prediction. When sampling with randomness, candidate acceptance depends on the probability distribution comparison between the target and draft models, with rejected tokens resampled from an adjusted distribution. The resulting token probabilities depend strictly on the target model, with deviations occurring only due to inherent sampling variance or numerical precision limits.
- Speculative decoding preserves the target model's exact output distribution through an acceptance rule.
- Under greedy decoding, a candidate token is retained only if it equals the target model's highest-probability choice.
- Under sampling, a candidate is always accepted if the target model assigns it equal or higher probability than the draft model, and probabilistically accepted if lower.
- Rejected candidates are replaced by sampling from an adjusted distribution that subtracts the draft model's probabilities, ensuring mathematical equivalence to the target model.
- Outputs can still differ between runs due to inherent sampling randomness and hardware-level floating-point rounding limits.
Acceptance Rate
In candidate-token acceleration approaches, the speedup is primarily dictated by the acceptance rate, which represents the fraction of proposed tokens preserved by the target model. Repetitive or structured tasks like code generation and summarization yield high acceptance rates, whereas open-ended tasks like creative writing result in frequent model divergence and lower acceptance. Higher sampling temperatures lower acceptance further by flattening output probability distributions, with the process becoming inefficient when acceptance drops below roughly 50 percent. Demonstrating this in practice, DeepSeek documented an 80 to 90 percent acceptance rate for the second predicted token during production serving of DeepSeek-V3, which yielded roughly 1.8x generation throughput.
- Acceptance rate is the fraction of candidate tokens kept by the target model, while acceptance length includes confirmed tokens plus one trailing free token.
- Structured and repetitive workloads (such as code generation, summarization, and retrieval-augmented answers) yield high acceptance rates.
- Open-ended outputs (such as creative writing and conversation) yield low acceptance rates due to small model divergence.
- Higher sampling temperature flattens token probability distributions and decreases acceptance.
- Acceptance rates below approximately 50% result in the computational overhead exceeding the efficiency gains.
- DeepSeek achieved 80% to 90% acceptance on the second predicted token in production serving of DeepSeek-V3, reaching roughly 1.8x generation throughput.
Candidate or Draft Sources
Speculative drafting relies on four primary mechanisms to obtain fast, low-cost token predictions: deploying a separate smaller sibling model, adding multi-token prediction heads to the target architecture, executing a lower-precision or pruned variant of the same weights, or retrieving matching sequences from prior context. Each option introduces distinct trade-offs between hardware overhead, training requirements, implementation complexity, and contextual applicability. Tokenizer alignment strictly dictates draft compatibility, often presenting a higher barrier than model quality differences. Ultimately, all four drafting strategies necessitate sufficient spare compute capacity during serving to provide a net speedup.
- Using a separate smaller model requires an identical tokenizer and family, consuming server VRAM from the KV cache budget.
- Lightweight extra prediction heads can be trained alongside the base model during pretraining and repurposed for inference drafting, as seen in DeepSeek-V3.
- QuantSpec executes draft generation using 4-bit weights and a 4-bit KV cache, yielding speedups above 1.78x with acceptance rates exceeding 90 percent.
- Context-based text searching requires zero memory overhead and provides 2x to 4x acceleration on repetitive tasks such as document editing and summarization.
- A draft model cannot be paired with a target model having a different vocabulary unless additional translation machinery is implemented.
Concurrency Limits
The speedup provided by speculative decoding depends significantly on server load and operating concurrency. As concurrent requests increase, compute units approach saturation and verification must compete with incoming requests, causing latency gains to diminish and potentially fall below baseline throughput. Serving systems such as vLLM manage this tradeoff by dynamically reducing draft length or disabling speculation beyond certain batch sizes. Additionally, because the technique targets token generation rather than prompt evaluation, it offers minimal latency benefits for workloads characterized by long prompts and short outputs.
- Speculative decoding gains diminish as server load increases and compute capacity becomes saturated.
- Empirical evaluations on a 70B model showed acceleration dropping from 1.96x at batch size 1 down to 1.21x at batch size 128, with potential net throughput losses under heavy concurrency.
- vLLM allows users to configure batch size thresholds to dynamically reduce draft length or disable speculation entirely.
- Time to first token remains unchanged by speculative decoding, limiting its utility for long-prompt, short-completion workloads.
- DeepSeek observed that multi-token prediction slightly reduces throughput while noticeably improving end-to-end generation latency.
Conclusion
Speculative decoding reorganizes the timing of token generation rather than reducing the total work executed by the target model. It leverages the transformer architecture's ability to evaluate multiple candidate tokens in parallel for roughly the cost of a single token. When drafted tokens are rejected, the verification pass still produces a correct token at the point of mismatch, preventing wasted computation. Consequently, baseline output quality is maintained by design through the acceptance rule without additional tuning.
- Token generation is memory-bandwidth bound because reading model weights dominates over arithmetic execution.
- Transformers can evaluate multiple candidate tokens in a single parallel pass for approximately the same cost as evaluating one.
- Draft rejections truncate rather than waste compute because the verification step yields a valid token at the mismatch position.
- Output quality is preserved directly by the acceptance rule without requiring tuning.
- Performance speedups depend on output predictability, server spare compute capacity, and draft generation costs.