← 回到 Reading
Daily Dose of DS 2026-08-24

Preloading Knowledge Into a Model Instead of Retrieving It

The RAG Systems course introduces preloading as a method to eliminate repetitive document processing by having the model read a knowledge base once and storing the resulting KV cache. This approach addresses the high cost of prefill, which scales quadratically with input length and often dominates inference expenses in production environments. While providers like Gemini and Claude offer significant discounts for cached tokens, technical challenges such as effective context limits and the need for query-agnostic compression must be addressed for successful implementation. The efficiency and cost of AI agents are primarily determined by the 'harness'—the runtime layer that manages context, tool execution, and memory—rather than the model itself. TrueFoundry's open-source harness, TrueForge, demonstrates how to reduce token usage by 2.7x through techniques like deferred tool loading, context offloading to sandboxes, and 'Code Mode' for data processing. Benchmarks on Enterprise-Bench show that these optimizations allow agents to achieve high success rates with significantly lower latency and cost compared to closed alternatives.

閱讀原文 ↗
目錄 2 段
  1. 01Preloading knowledge into a model instead of retrieving it
  2. 02How to cut agent tokens by 2.7x (using an open harness)
DEEP DIVE

Preloading knowledge into a model instead of retrieving it

The RAG Systems course introduces preloading as a method to eliminate repetitive document processing by having the model read a knowledge base once and storing the resulting KV cache. This approach addresses the high cost of prefill, which scales quadratically with input length and often dominates inference expenses in production environments. While providers like Gemini and Claude offer significant discounts for cached tokens, technical challenges such as effective context limits and the need for query-agnostic compression must be addressed for successful implementation.

  • Preloading allows a model to skip retrieval, chunking, and embedding by storing the KV cache of a corpus.
  • Prefill costs scale quadratically with input length, making it the dominant expense for high-volume RAG systems.
  • Provider-side caching for models like Gemini and Claude can reduce input token costs by up to 90 percent.
  • A model's effective context window is often significantly smaller than its advertised limit, with performance degrading as context grows.
  • Most KV cache compression methods are unsuitable for preloading because they require the query at the time of compression.
  • Modular preloading caches passages independently but can encounter cross-attention failures.
  • Trained preloading involves distilling a corpus into a compact cache through a specific training run.
HARNESS ENGINEERING

How to cut agent tokens by 2.7x (using an open harness)

The efficiency and cost of AI agents are primarily determined by the 'harness'—the runtime layer that manages context, tool execution, and memory—rather than the model itself. TrueFoundry's open-source harness, TrueForge, demonstrates how to reduce token usage by 2.7x through techniques like deferred tool loading, context offloading to sandboxes, and 'Code Mode' for data processing. Benchmarks on Enterprise-Bench show that these optimizations allow agents to achieve high success rates with significantly lower latency and cost compared to closed alternatives.

  • Agent token costs are often a runtime problem caused by redundant context and inefficient tool definitions being sent in every model call.
  • TrueForge achieved a 2.7x reduction in token usage compared to Claude Managed Agents while maintaining the same task success rate.
  • Deferred tool loading prevents prompt bloat by only loading full schemas and documentation when a specific tool is identified as necessary.
  • Code Mode reduces context usage by executing Python scripts in a sandbox to join and process data, returning only the final result to the model.
  • Context offloading moves large tool responses into a sandbox file, providing the model with a short preview and file path instead of the full payload.
  • TrueForge is MIT-licensed, self-hostable, and integrates with the Model Context Protocol (MCP) and Daytona sandboxes.