← 回到 Reading
Daily Dose of DS 2026-09-14

What It Takes to Build a Production Agent Harness

An agent harness serves as the execution layer that manages interactions between models, tools, and application state. This series demonstrates building such a harness using LangChain and LangGraph to transition from simple requests to stateful production applications. The implementation addresses critical system design challenges including error propagation, state persistence, and deterministic execution control. The section describes a self-repairing agent harness system using Hermes and Opik to learn from production failures without retraining model weights. It distinguishes between runtime learning, where successful procedures are saved as skills, and offline optimization using the GEPA algorithm for prompt and skill evolution. The workflow involves capturing execution traces, diagnosing root causes with the Ollie tool, and establishing regression tests to prevent recurring errors. By modifying external artifacts like instructions and configurations, the system converts production incidents into structured, reviewable inputs for continuous improvement.

閱讀原文 ↗
目錄 2 段
  1. 01Building a production agent harness in LangChain
  2. 02Your agent harness should repair itself
AGENT ENGINEERING

Building a production agent harness in LangChain

An agent harness serves as the execution layer that manages interactions between models, tools, and application state. This series demonstrates building such a harness using LangChain and LangGraph to transition from simple requests to stateful production applications. The implementation addresses critical system design challenges including error propagation, state persistence, and deterministic execution control.

  • An agent harness is the execution layer responsible for managing the interaction between models, tools, and state.
  • Production agent platforms require capabilities like sessions, checkpoints, tracing, and human-in-the-loop approval.
  • LangChain provides interfaces for models, messages, and tools, while LangGraph offers explicit state transitions and execution control.
  • System design for agents involves deciding which state is run-specific versus persistent across conversations.
  • The harness must handle tool failures by deciding whether to retry, request human input, or stop execution.
HANDS-ON

Your agent harness should repair itself

The section describes a self-repairing agent harness system using Hermes and Opik to learn from production failures without retraining model weights. It distinguishes between runtime learning, where successful procedures are saved as skills, and offline optimization using the GEPA algorithm for prompt and skill evolution. The workflow involves capturing execution traces, diagnosing root causes with the Ollie tool, and establishing regression tests to prevent recurring errors. By modifying external artifacts like instructions and configurations, the system converts production incidents into structured, reviewable inputs for continuous improvement.

  • Agent systems can improve by modifying external artifacts like instructions, skills, and configurations instead of updating model weights.
  • Hermes preserves successful task procedures as editable SKILL.md files, which serve as procedural memory for future reuse.
  • The GEPA (Genetic-Pareto Prompt Evolution) method enables offline optimization of agent artifacts without requiring GPU fine-tuning.
  • Opik provides a tracing infrastructure that captures model spans and tool spans to pinpoint the exact step where a multi-step decision process failed.
  • Ollie acts as a diagnostic bridge that inspects project files and proposes code or configuration diffs based on trace data.
  • A robust repair loop requires human approval of proposed changes followed by verification through reruns and the creation of regression test cases.
  • Opik Test Suites use natural-language assertions and LLM judges to validate agent behavior against expected outcomes.