← 回到 Reading
Daily Dose of DS 2026-08-17

How a GPU Actually Works

AI application development often incurs high costs due to real API calls during continuous integration (CI) cycles. CopilotKit has introduced aimock, an open-source tool that provides a local mock server to replace these expensive calls. Unlike static mocks, aimock maintains schema accuracy by daily testing its responses against actual provider APIs and official client libraries. This approach allows developers to run offline tests while ensuring their integrations remain compatible with the latest provider updates across multiple platforms. GPU performance during LLM inference is primarily constrained by memory bandwidth rather than raw computational power. While arithmetic operations are inexpensive, fetching data from memory is costly, leading to a 'memory-bound' state where arithmetic units remain idle. Current hardware requires approximately 300 operations per byte to reach peak efficiency, yet token generation typically performs only two operations per weight. Consequently, common optimization techniques like batching and quantization are designed to either increase the work done per fetch or reduce the total data transferred. Prefix caching is a powerful optimization for stable prompts, offering up to 90% cost reduction, but it fails when document order changes or multiple documents are combined due to its strict byte-for-byte prefix requirement. CacheBlend, a new method from the LMCache project, solves this by allowing modular reuse of cached document states and only recomputing boundary tokens. This approach enables 2-4x faster processing for multi-document queries without quality loss. LMCache serves as an open-source management layer that integrates this technology with major inference engines like vLLM and SGLang.

閱讀原文 ↗
目錄 3 段
  1. 01Mock infrastructure for AI apps
  2. 02How a GPU actually works
  3. 03Prefix caching vs CacheBlend
OPEN-SOURCE

Mock infrastructure for AI apps

AI application development often incurs high costs due to real API calls during continuous integration (CI) cycles. CopilotKit has introduced aimock, an open-source tool that provides a local mock server to replace these expensive calls. Unlike static mocks, aimock maintains schema accuracy by daily testing its responses against actual provider APIs and official client libraries. This approach allows developers to run offline tests while ensuring their integrations remain compatible with the latest provider updates across multiple platforms.

  • CI runs for AI applications frequently incur significant costs by making real requests to LLM providers.
  • Traditional manual mocking often fails when providers update their API schemas without notice.
  • aimock is an open-source project by CopilotKit that serves as a local mock server for various AI APIs.
  • The tool ensures schema validity by running daily comparisons between its mock responses and real API outputs.
  • When schema changes are detected, aimock updates its built-in definitions and releases a patch via npm.
  • aimock supports a wide range of providers and tools, including OpenAI, Claude, Gemini, and vector databases like Pinecone.
DEEP DIVE

How a GPU actually works

GPU performance during LLM inference is primarily constrained by memory bandwidth rather than raw computational power. While arithmetic operations are inexpensive, fetching data from memory is costly, leading to a 'memory-bound' state where arithmetic units remain idle. Current hardware requires approximately 300 operations per byte to reach peak efficiency, yet token generation typically performs only two operations per weight. Consequently, common optimization techniques like batching and quantization are designed to either increase the work done per fetch or reduce the total data transferred.

  • GPU utilization metrics can be misleading because a chip waiting for data appears identical to one performing active computation.
  • The break-even point for modern hardware efficiency is roughly 300 operations per byte at 16-bit precision.
  • Token generation for a 70B model is heavily memory-bound, achieving only about 24 tokens per second on hardware with 3.3 TB/s bandwidth.
  • The GPU memory hierarchy consists of the register file, shared memory, L2 cache, and HBM.
  • Optimization strategies like FlashAttention, fusion, and quantization all focus on improving the ratio of arithmetic operations to memory fetches.
LLMOPS

Prefix caching vs CacheBlend

Prefix caching is a powerful optimization for stable prompts, offering up to 90% cost reduction, but it fails when document order changes or multiple documents are combined due to its strict byte-for-byte prefix requirement. CacheBlend, a new method from the LMCache project, solves this by allowing modular reuse of cached document states and only recomputing boundary tokens. This approach enables 2-4x faster processing for multi-document queries without quality loss. LMCache serves as an open-source management layer that integrates this technology with major inference engines like vLLM and SGLang.

  • Prefix caching requires an exact byte-for-byte match, making it fragile for RAG and dynamic conversation histories.
  • Alibaba Cloud data indicates that 10% of KV cache blocks account for 77% of all cache hits.
  • CacheBlend allows for the reuse of document caches regardless of their order or surrounding context by selectively recomputing boundary tokens.
  • CacheBlend provides a 2x to 4x speedup for multi-document queries compared to standard prefix caching.
  • LMCache is an open-source cache management layer that operates outside the inference engine to support vLLM, SGLang, and TensorRT-LLM.
  • The CacheBlend research paper received the Best Paper Award at EuroSys 2025.