← 回到 Reading
Daily Dose of DS 2026-07-16

Agents Need a New Kind of Web Search

AI-generated code and prompt changes introduce 'unknown-unknown' failure modes that traditional testing and staging environments often fail to capture. Because these systems lack a human mental model, production becomes the primary validation environment, necessitating high-cardinality telemetry for effective debugging. The updated 'Observability Engineering' book addresses these shifts with new content on instrumenting LLM applications and leveraging agentic AI for system maintenance. AI agents currently face a "retrieval tax," consuming excessive tokens to fetch and clean raw web content from search snippets. Using an "owned index" like Seltz addresses this by providing pre-processed, structured documents directly via API, which can be up to four times more cost-effective than traditional search loops. This structured approach allows agents to focus on reasoning and perform complex data joins across multiple records. Developers are encouraged to chain discovery-based search with document-based retrieval to optimize both performance and cost. The evolution of AI retrieval systems is moving from static RAG to Agentic RAG and finally to AI Memory. While RAG and Agentic RAG are primarily read-only processes, AI Memory introduces read-write capabilities that allow agents to learn from past interactions and user preferences. This shift enables continual learning without the need for model retraining, though it introduces new challenges like memory corruption and management of different memory types.

閱讀原文 ↗
目錄 4 段
  1. 01A free O’Reilly book on debugging production systems
  2. 02Agents need a new kind of web search
  3. 03RAG, Agentic RAG, and AI Memory
  4. 04Knowledge Distillation using Teacher Assistant
TOGETHER WITH HONEYCOMB

A free O’Reilly book on debugging production systems

AI-generated code and prompt changes introduce 'unknown-unknown' failure modes that traditional testing and staging environments often fail to capture. Because these systems lack a human mental model, production becomes the primary validation environment, necessitating high-cardinality telemetry for effective debugging. The updated 'Observability Engineering' book addresses these shifts with new content on instrumenting LLM applications and leveraging agentic AI for system maintenance.

  • Coding agents reduce the cost of code creation but do not improve validation capacity.
  • AI-generated code is prone to 'unknown-unknown' failures because no human mental model is formed during its creation.
  • Offline evaluations of prompt changes often fail to predict production degradation due to traffic mismatches.
  • Effective production debugging requires raw, high-cardinality event data that can be filtered by attributes like user ID or prompt version.
  • The 'Observability Engineering' book has been significantly updated with 27 new chapters on LLM apps and agentic AI debugging.
AGENTS

Agents need a new kind of web search

AI agents currently face a "retrieval tax," consuming excessive tokens to fetch and clean raw web content from search snippets. Using an "owned index" like Seltz addresses this by providing pre-processed, structured documents directly via API, which can be up to four times more cost-effective than traditional search loops. This structured approach allows agents to focus on reasoning and perform complex data joins across multiple records. Developers are encouraged to chain discovery-based search with document-based retrieval to optimize both performance and cost.

  • The "retrieval tax" refers to tokens spent on fetching and cleaning web content before reasoning can occur.
  • Traditional search APIs return snippets and links, forcing agents to perform manual scraping and HTML cleaning.
  • An owned index like Seltz provides full, structured documents in a single API call.
  • Multi-hop web search loops can cost 4x more in tokens than using a pre-indexed document service.
  • Seltz offers specialized scopes for People, News, and Wikipedia content to provide structured context.
  • Full document access enables agents to answer questions that require intersecting data from multiple sources.
LLMs

RAG, Agentic RAG, and AI Memory

The evolution of AI retrieval systems is moving from static RAG to Agentic RAG and finally to AI Memory. While RAG and Agentic RAG are primarily read-only processes, AI Memory introduces read-write capabilities that allow agents to learn from past interactions and user preferences. This shift enables continual learning without the need for model retraining, though it introduces new challenges like memory corruption and management of different memory types.

  • RAG systems (2020-2023) are limited to one-shot retrieval without decision-making capabilities.
  • Agentic RAG allows agents to decide when and where to retrieve information, though it remains read-only.
  • AI Memory enables agents to both read from and write to external knowledge stores, facilitating true personalization.
  • Continual learning allows AI agents to accumulate knowledge over time without requiring retraining.
  • AI Memory management involves handling procedural, episodic, and semantic memory types.
  • Graphiti is an open-source framework designed to help build real-time knowledge graphs for agent memory.
MACHINE LEARNING

Knowledge Distillation using Teacher Assistant

Knowledge distillation is a technique used to compress large machine learning models by training a smaller student model to mimic a larger teacher model. However, a significant size gap between the teacher and student can lead to performance degradation, as students have a limit on the teacher size they can effectively learn from. The Teacher Assistant Knowledge Distillation (TAKD) method addresses this by introducing an intermediate assistant model to bridge the gap. This multi-step process involves the assistant learning from the teacher and the student subsequently learning from the assistant, resulting in improved accuracy over direct distillation.

  • Student model accuracy can decrease if the teacher model is too large relative to the student's capacity.
  • Knowledge distillation is most effective when the size difference between models is within a specific range.
  • The Teacher Assistant Knowledge Distillation (TAKD) method uses an intermediate model to bridge the gap between a large teacher and a small student.
  • TAKD involves a two-step training process: the assistant model learns from the teacher, then the student learns from the assistant.
  • Experimental results show TAKD outperforms both direct distillation (BLKD) and training without distillation (NOKD).
  • The assistant model can be significantly smaller, such as 50% smaller, than the teacher model.
  • While TAKD adds a training step, it enhances production performance and efficiency where costs are typically higher.