Easiest Way to Run Agent Harnesses Using Local Models
Dynatrace has released a reference application to help developers inspect and debug the internal operations of LLM pipelines. The application demonstrates how to use distributed tracing to measure latency across embeddings, retrieval, and generation steps rather than treating the pipeline as a single opaque call. Built with Python, it utilizes Amazon Bedrock for core AI tasks and OpenTelemetry for observability, providing a template for instrumenting production LLM apps. Magnitude is an open-source tool designed to simplify the process of running local LLMs for agent harnesses. It addresses the complexity of model selection by profiling hardware and benchmarking performance rather than just checking memory capacity. The tool recommends practical models based on context length and quantization, then facilitates connection to agents like Claude Code or OpenCode. This approach ensures that the selected model is fast enough for the repetitive workloads typical of agent loops. Prompt injection is identified by OWASP as the primary threat to LLM applications, necessitating a shift from behavioral constraints to architectural enforcement. Five key defenses are proposed: using delimiters for untrusted text, establishing an instruction hierarchy, applying the principle of least privilege to tools, requiring human-in-the-loop approval, and implementing the Dual LLM pattern. These strategies, such as those formalized in Google's Spotlight and DeepMind's CaMeL framework, aim to create robust security boundaries by layering multiple protective measures.
閱讀原文 ↗目錄
Tracing what happens inside an LLM pipeline
Dynatrace has released a reference application to help developers inspect and debug the internal operations of LLM pipelines. The application demonstrates how to use distributed tracing to measure latency across embeddings, retrieval, and generation steps rather than treating the pipeline as a single opaque call. Built with Python, it utilizes Amazon Bedrock for core AI tasks and OpenTelemetry for observability, providing a template for instrumenting production LLM apps.
- Measuring only the outer HTTP request latency is insufficient for debugging complex LLM pipelines involving retrieval and agents.
- Dynatrace's reference application includes a Python travel advisor with basic, RAG, and agentic execution paths.
- The system uses Amazon Bedrock for both text generation and embedding tasks.
- Observability is achieved using OpenTelemetry and Traceloop to record workflows and export traces via OTLP.
- The reference repository provides sample dashboards and deployment files for local or Kubernetes-based execution.
Easiest way to run agent harnesses using local models
Magnitude is an open-source tool designed to simplify the process of running local LLMs for agent harnesses. It addresses the complexity of model selection by profiling hardware and benchmarking performance rather than just checking memory capacity. The tool recommends practical models based on context length and quantization, then facilitates connection to agents like Claude Code or OpenCode. This approach ensures that the selected model is fast enough for the repetitive workloads typical of agent loops.
- Local model selection depends on weights, context length (KV-cache), and quantization levels.
- Available RAM or VRAM is an insufficient metric for determining if inference speed is adequate for agent loops.
- Agent harnesses generate long, repeated workloads that differ from isolated prompt-response cycles.
- Magnitude profiles hardware and benchmarks performance to recommend practical local models.
- Magnitude supports integration with harnesses such as Claude Code, Codex, OpenCode, and Pi.
Practical defenses for prompt injection
Prompt injection is identified by OWASP as the primary threat to LLM applications, necessitating a shift from behavioral constraints to architectural enforcement. Five key defenses are proposed: using delimiters for untrusted text, establishing an instruction hierarchy, applying the principle of least privilege to tools, requiring human-in-the-loop approval, and implementing the Dual LLM pattern. These strategies, such as those formalized in Google's Spotlight and DeepMind's CaMeL framework, aim to create robust security boundaries by layering multiple protective measures.
- OWASP ranks prompt injection as the number one threat to LLM-based applications.
- Delimiters like XML tags or Base64 encoding provide structural signals to models to treat input as data rather than instructions.
- Instruction hierarchy assigns trust levels to different prompt sources, ensuring developer instructions override third-party content.
- The principle of least privilege reduces risk by restricting an agent's tool access to the absolute minimum required for a task.
- The Dual LLM pattern separates reasoning (Planner) from action (Executor) to prevent untrusted data from influencing control flow.
- Google DeepMind's CaMeL framework successfully addressed the AgentDojo security benchmark using architectural isolation.