← 回到 Reading
Daily Dose of DS 2026-08-27

KV vs Prefix vs Prompt vs Semantic Caching

Traditional graph databases like Neo4j are designed for single large graphs, making them inefficient for agent memory workloads that require millions of small, per-user graphs. Zep developed Konig to solve this by implementing a tiered storage architecture that moves idle graphs to object storage while keeping active ones in RAM. This design allows for per-graph encryption, sub-200ms end-to-end latency, and inline execution of algorithms like PageRank. Konig also leverages AVX-512 search kernels and bi-temporal facts to optimize performance and data management. The article distinguishes between four types of caching in LLM stacks: KV, Prefix, Prompt, and Semantic caching. While the first three are exact-match techniques that store attention tensors to optimize prefill and decoding performance, Semantic caching uses fuzzy-match embeddings to return stored responses. The text details how these caches impact memory bandwidth and cost, while also identifying common failure modes like dynamic prompt headers that invalidate cache blocks. Technical optimizations such as Grouped-query attention and FP8 quantization are highlighted as methods to manage the high memory requirements of these systems. Temperature is a hyperparameter in Large Language Models (LLMs) that adjusts the probability distribution of the next-token prediction by modifying the softmax function. Low temperature values concentrate probability on the most likely tokens, making the model's output more deterministic and predictable. High temperature values spread the probability more evenly across the vocabulary, resulting in more diverse, creative, or potentially nonsensical responses.

閱讀原文 ↗
目錄 3 段
  1. 01Why per-user memory graphs break normal graph DBs
  2. 02KV vs Prefix vs Prompt vs Semantic Caching
  3. 03What is Temperature in LLMs?
AGENT MEMORY

Why per-user memory graphs break normal graph DBs

Traditional graph databases like Neo4j are designed for single large graphs, making them inefficient for agent memory workloads that require millions of small, per-user graphs. Zep developed Konig to solve this by implementing a tiered storage architecture that moves idle graphs to object storage while keeping active ones in RAM. This design allows for per-graph encryption, sub-200ms end-to-end latency, and inline execution of algorithms like PageRank. Konig also leverages AVX-512 search kernels and bi-temporal facts to optimize performance and data management.

  • Traditional graph databases like Neo4j are optimized for single large graphs, whereas agent memory requires millions of small, isolated graphs.
  • Konig manages millions of graphs by using a tiered storage system involving RAM, local NVMe, and object storage.
  • Per-graph isolation in Konig allows for individual encryption keys and retention rules for each user or project.
  • Konig enables inline PageRank execution in milliseconds by limiting query scope to a single graph.
  • The system maintains sub-200ms end-to-end latency even when scaling from thousands to tens of millions of graphs.
  • Konig utilizes AVX-512 search kernels and supports bi-temporal facts for advanced data processing.
LLMs

KV vs Prefix vs Prompt vs Semantic Caching

The article distinguishes between four types of caching in LLM stacks: KV, Prefix, Prompt, and Semantic caching. While the first three are exact-match techniques that store attention tensors to optimize prefill and decoding performance, Semantic caching uses fuzzy-match embeddings to return stored responses. The text details how these caches impact memory bandwidth and cost, while also identifying common failure modes like dynamic prompt headers that invalidate cache blocks. Technical optimizations such as Grouped-query attention and FP8 quantization are highlighted as methods to manage the high memory requirements of these systems.

  • KV caching stores key and value vectors to avoid recomputing the full sequence during decoding, shifting the bottleneck from compute to memory bandwidth.
  • Prefix caching, used by engines like vLLM, stores KV tensors in fixed-size blocks identified by hash chains to allow sharing across requests.
  • Prompt caching is a commercial version of prefix caching offered by providers like Anthropic and OpenAI with specific read/write pricing.
  • Semantic caching returns stored response strings based on embedding similarity, which can lead to incorrect answers if the similarity threshold is too low.
  • Dynamic content like timestamps or user IDs at the start of a prompt will invalidate all subsequent cache blocks.
  • Advanced architectures like Multi-head latent attention (MLA) in DeepSeek models compress KV caches into latent vectors to save space.
LLMs

What is Temperature in LLMs?

Temperature is a hyperparameter in Large Language Models (LLMs) that adjusts the probability distribution of the next-token prediction by modifying the softmax function. Low temperature values concentrate probability on the most likely tokens, making the model's output more deterministic and predictable. High temperature values spread the probability more evenly across the vocabulary, resulting in more diverse, creative, or potentially nonsensical responses.

  • Temperature modifies the softmax function applied to logits to influence token sampling.
  • Low temperature values create a peaked distribution, leading to greedy and predictable token selection.
  • High temperature values create a more uniform distribution, resulting in stochastic and creative outputs.
  • Traditional classification models are typically deterministic, whereas LLMs rely on sampling from probabilities.
  • Excessive temperature values can degrade output quality into gibberish.