← 回到 Reading
Daily Dose of DS 2026-08-30

Why KV Cache Stores K and V Vectors But Never Q?

Datalab Marker v2 is an open-source document parsing pipeline designed to convert PDFs, images, and office documents into clean Markdown, JSON, or HTML. It utilizes a shared inference server architecture to maximize GPU throughput by batching requests from multiple CPU workers, achieving up to 23.7 pages per second on a B200 GPU. The system is powered by Surya 2, a 650M parameter model that handles OCR, layout analysis, and table recognition across 90+ languages. LLMs operate autoregressively, where each new token is generated using only the last hidden state of the previous pass. In the attention mechanism, a query vector (Q) is used only once to compute the attention row for its specific position, whereas key (K) and value (V) vectors are reused in every subsequent decoding step. Because K and V vectors remain constant for a given token due to causal masking, they are stored in a KV cache to avoid redundant computations, while Q vectors are discarded after their single use. This section provides a guide on building a Model Context Protocol (MCP) server for multimodal AI applications. The architecture utilizes Pixeltable for infrastructure and data management across various modalities like video and audio, while CrewAI handles the orchestration of the agentic workflow. The system design features a router agent that directs queries to specialized agents to generate coherent responses based on the input modality.

閱讀原文 ↗
目錄 3 段
  1. 01Turn any PDF, image, DOCX, and PPTX into clean MD
  2. 02Why KV cache stores K and V vectors but never Q?
  3. 03Build the ultimate MCP Server for multimodal AI
OPEN-SOURCE

Turn any PDF, image, DOCX, and PPTX into clean MD

Datalab Marker v2 is an open-source document parsing pipeline designed to convert PDFs, images, and office documents into clean Markdown, JSON, or HTML. It utilizes a shared inference server architecture to maximize GPU throughput by batching requests from multiple CPU workers, achieving up to 23.7 pages per second on a B200 GPU. The system is powered by Surya 2, a 650M parameter model that handles OCR, layout analysis, and table recognition across 90+ languages.

  • Marker v2 achieves high throughput by using a shared Surya inference server that batches requests from multiple lightweight CPU workers.
  • The underlying Surya 2 model is a single 650M parameter model performing OCR, layout, reading order, and table recognition.
  • Three operational modes are available: balanced (high quality), fast (text-layer focused), and disable_ocr (CPU-only).
  • On the olmOCR-bench, Marker v2 scored 76.0%, outperforming MinerU (72.7%) and Docling (50.3%) while maintaining higher throughput.
  • Fast mode is significantly faster but less accurate for math-heavy documents as it relies on the PDF's native text layer.
  • The pipeline supports over 90 languages and handles complex elements like equations and low-confidence tables.
LLMs

Why KV cache stores K and V vectors but never Q?

LLMs operate autoregressively, where each new token is generated using only the last hidden state of the previous pass. In the attention mechanism, a query vector (Q) is used only once to compute the attention row for its specific position, whereas key (K) and value (V) vectors are reused in every subsequent decoding step. Because K and V vectors remain constant for a given token due to causal masking, they are stored in a KV cache to avoid redundant computations, while Q vectors are discarded after their single use.

  • The prefill stage processes the entire prompt in parallel and is the primary contributor to Time to First Token (TTFT).
  • During decoding, only the query vector of the current token is required to compute the attention output for that step.
  • Key and value vectors are reused across all future decoding steps because causal masking prevents future tokens from affecting past K and V values.
  • Query vectors are position-specific and are never reused once the hidden state for that position is generated.
  • The KV cache is one of four common caching layers, alongside prefix caching, prompt caching, and semantic caching.
HANDS-ON

Build the ultimate MCP Server for multimodal AI

This section provides a guide on building a Model Context Protocol (MCP) server for multimodal AI applications. The architecture utilizes Pixeltable for infrastructure and data management across various modalities like video and audio, while CrewAI handles the orchestration of the agentic workflow. The system design features a router agent that directs queries to specialized agents to generate coherent responses based on the input modality.

  • Pixeltable is an open-source Python library designed to streamline multimodal AI pipelines from storage to execution.
  • CrewAI is used to orchestrate the agentic workflow within the multimodal system.
  • The system architecture includes a router agent that identifies the modality of a user query and triggers a specialist agent.
  • Pixeltable supports multiple data types including images, videos, text, and audio.
  • The MCP server is built directly on top of the Pixeltable infrastructure.