← 回到 Reading
Daily Dose of DS 2026-09-07

LLM Routing Can Cost More Than Not Routing

Honeycomb is hosting a free six-session live masterclass on observability engineering led by Liz Fong-Jones, co-author of the O'Reilly book on the subject. The course addresses the challenges of debugging non-deterministic systems like LLM agents by emphasizing context-rich request recording over traditional dashboards. It covers practical topics such as OpenTelemetry instrumentation, cost optimization, and monitoring production deployments for mobile and ML systems. The sessions include hands-on labs and live Q&A, with recordings available for those who missed earlier dates. LLM routing optimizes costs by directing tasks to the most efficient model, but traditional implementations often suffer from high overhead and cache invalidation. DigitalOcean's Inference Router addresses these challenges by moving routing logic into the infrastructure layer using specialized, small-parameter models for intent resolution. Key features include session pinning to preserve prefix caching and real-time model ranking based on latency and price. This approach allows agentic workflows, which consume significantly more tokens, to remain cost-effective without sacrificing performance. Red teaming has become a critical component of LLM development, with major labs like OpenAI, Meta, and Google investing heavily to identify vulnerabilities that standard accuracy metrics miss. While manual red teaming is resource-intensive, the open-source framework DeepTeam automates this process by simulating over 10 attack methods and detecting more than 40 types of vulnerabilities. The tool supports both single-turn and multi-turn conversational testing and generates detailed risk assessments without requiring pre-existing datasets. DeepTeam also integrates with platforms like Confident AI for comprehensive risk logging and dashboarding.

閱讀原文 ↗
目錄 3 段
  1. 01Free Observability Engineering Masterclass with Liz Fong-Jones & Honeycomb
  2. 02LLM routing can cost more than not routing
  3. 03Hands-on guide to Red Teaming LLM apps
FREE MASTERCLASS

Free Observability Engineering Masterclass with Liz Fong-Jones & Honeycomb

Honeycomb is hosting a free six-session live masterclass on observability engineering led by Liz Fong-Jones, co-author of the O'Reilly book on the subject. The course addresses the challenges of debugging non-deterministic systems like LLM agents by emphasizing context-rich request recording over traditional dashboards. It covers practical topics such as OpenTelemetry instrumentation, cost optimization, and monitoring production deployments for mobile and ML systems. The sessions include hands-on labs and live Q&A, with recordings available for those who missed earlier dates.

  • Traditional dashboards are often ineffective for debugging non-deterministic LLM agent failures because they are built around known failure modes.
  • Effective observability involves recording enough context per request to allow for post-hoc analysis on unplanned variables.
  • Liz Fong-Jones, co-author of the O'Reilly book Observability Engineering, is the instructor for the masterclass.
  • The masterclass includes practical training on using OpenTelemetry for system instrumentation.
  • Upcoming sessions focus on performance thresholds, cost management, and monitoring mobile apps and ML systems.
  • Honeycomb provides recordings and hands-on labs for all sessions in the series.
HANDS-ON

LLM routing can cost more than not routing

LLM routing optimizes costs by directing tasks to the most efficient model, but traditional implementations often suffer from high overhead and cache invalidation. DigitalOcean's Inference Router addresses these challenges by moving routing logic into the infrastructure layer using specialized, small-parameter models for intent resolution. Key features include session pinning to preserve prefix caching and real-time model ranking based on latency and price. This approach allows agentic workflows, which consume significantly more tokens, to remain cost-effective without sacrificing performance.

  • Naive routing can be more expensive than single-model usage due to classification costs and prefix cache invalidation.
  • Specialized small models like the 1.5B parameter Arch-Router can outperform frontier models in routing accuracy and speed.
  • Session pinning via the X-Model-Affinity header allows agent loops to reuse attention state caches across multiple turns.
  • The Inference Router uses real-time metrics from Prometheus and pricing APIs to dynamically rank model pools.
  • Agentic workflows are estimated by Gartner to use 5 to 30 times more tokens than standard chat interactions.
  • DigitalOcean provides pre-configured routers for specific workloads like software engineering and document intelligence.
LLMs

Hands-on guide to Red Teaming LLM apps

Red teaming has become a critical component of LLM development, with major labs like OpenAI, Meta, and Google investing heavily to identify vulnerabilities that standard accuracy metrics miss. While manual red teaming is resource-intensive, the open-source framework DeepTeam automates this process by simulating over 10 attack methods and detecting more than 40 types of vulnerabilities. The tool supports both single-turn and multi-turn conversational testing and generates detailed risk assessments without requiring pre-existing datasets. DeepTeam also integrates with platforms like Confident AI for comprehensive risk logging and dashboarding.

  • OpenAI launched a $500,000 red-teaming challenge in August 2025 specifically for the gpt-oss-20b model.
  • Standard evaluation metrics like correctness and faithfulness fail to capture how easily a model can be exploited through prompt injections or jailbreaking.
  • DeepTeam is an open-source framework that enables end-to-end LLM red teaming with just a few lines of code.
  • The framework can detect vulnerabilities including PII leakage, bias, toxicity, and unauthorized access across 40+ categories.
  • DeepTeam differentiates between single-turn tests for immediate jailbreaks and multi-turn tests for conversational grooming.
  • Adversarial attacks in DeepTeam are dynamically simulated at runtime, removing the need for manual dataset creation.
  • Risk reports generated by DeepTeam can be logged and assessed via the Confident AI dashboard.