← 回到 Reading
Daily Dose of DS 2026-07-07

Rethinking KV Caching For Production Inference

Gitar, an AI-native code review tool now part of Sonar, automates the detection and fixing of bugs in pull requests by using full codebase context. It generates patches and validates them against CI systems to ensure builds pass before human intervention. This workflow, termed the Agent Centric Development Cycle (AC/DC), utilizes the Sonar Vortex engine for verification. The implementation of this cycle has been shown to reduce AI-related outages by 44% and token usage by 36%. Research from Stanford highlights that AI agents waste significant inference budgets on redundant content, with approximately 62% of data being repeated prompts and documents. While traditional prefix caching offers some cost reduction, it fails during multi-document RAG or when document order changes. LMCache addresses these limitations by decoupling cache management from the inference engine into a separate process, utilizing shared GPU memory to prevent performance bottlenecks. This architecture, paired with the CacheBlend technique, enables faster multi-document queries and significantly improves time-to-first-token performance. Large Language Models generate text sequentially by predicting the next token from a probability distribution rather than planning full sentences in advance. To optimize output, developers use various decoding strategies such as greedy search, random sampling, beam search, and contrastive methods. Each strategy offers different trade-offs between repetition, creativity, and global sequence probability. While some methods like beam search are ideal for accuracy-focused tasks like translation, others focus on maintaining diversity and coherence in long-form generation.

閱讀原文 ↗
目錄 3 段
  1. 01The pull request now fixes itself before anyone reads it
  2. 02Rethinking KV caching for production inference
  3. 034 LLM text generation strategies
TOGETHER WITH SONAR

The pull request now fixes itself before anyone reads it

Gitar, an AI-native code review tool now part of Sonar, automates the detection and fixing of bugs in pull requests by using full codebase context. It generates patches and validates them against CI systems to ensure builds pass before human intervention. This workflow, termed the Agent Centric Development Cycle (AC/DC), utilizes the Sonar Vortex engine for verification. The implementation of this cycle has been shown to reduce AI-related outages by 44% and token usage by 36%.

  • Gitar analyzes pull requests with full codebase context rather than just diffs to identify bugs.
  • The tool automatically generates patches and verifies them against CI before human review.
  • The Agent Centric Development Cycle (AC/DC) is a loop where an agent writes code and the Sonar Vortex engine verifies it.
  • Using Gitar and Sonar results in 44% fewer outages tied to AI-generated code.
  • Token usage is reduced by up to 36% due to decreased codebase re-parsing.
DEEP DIVE

Rethinking KV caching for production inference

Research from Stanford highlights that AI agents waste significant inference budgets on redundant content, with approximately 62% of data being repeated prompts and documents. While traditional prefix caching offers some cost reduction, it fails during multi-document RAG or when document order changes. LMCache addresses these limitations by decoupling cache management from the inference engine into a separate process, utilizing shared GPU memory to prevent performance bottlenecks. This architecture, paired with the CacheBlend technique, enables faster multi-document queries and significantly improves time-to-first-token performance.

  • Approximately 62% of content sent to AI agents per call consists of repeated system prompts, tool definitions, and documents.
  • Agentic workflows consume 5 to 30 times more tokens per task than standard chatbots because context is re-sent at every step.
  • Prefix caching requires an exact byte-for-byte match, leading to cache misses when document order changes or multiple documents are combined.
  • LMCache runs as a separate process from the inference engine to avoid resource contention between I/O-heavy cache management and compute-heavy inference.
  • CacheBlend, which won the EuroSys 2025 Best Paper Award, allows for 2-4x faster multi-document queries by selectively recomputing only tokens with cross-document connections.
  • LMCache provides production-grade features including Kubernetes operators, Prometheus integration, and support for vLLM, SGLang, and TensorRT-LLM.
LLMs

4 LLM text generation strategies

Large Language Models generate text sequentially by predicting the next token from a probability distribution rather than planning full sentences in advance. To optimize output, developers use various decoding strategies such as greedy search, random sampling, beam search, and contrastive methods. Each strategy offers different trade-offs between repetition, creativity, and global sequence probability. While some methods like beam search are ideal for accuracy-focused tasks like translation, others focus on maintaining diversity and coherence in long-form generation.

  • LLMs predict text token-by-token based on a probability vector for the next step.
  • Greedy search selects the highest probability token but frequently results in repetitive sentences.
  • Random sampling uses a temperature parameter to control the level of randomness in the generated output.
  • Beam search approximates global probability maximization by maintaining and expanding the top k candidate sequences.
  • Beam search is widely utilized in machine translation where correctness is more critical than creativity.
  • A newer contrastive search method penalizes tokens similar to previous context to prevent repetitive loops.
  • Contrastive search is particularly effective for maintaining coherence in long-form storytelling.