How to Reduce LLM Costs by 50-60% Using Model Routing
Model routing reduces LLM costs by automatically selecting the most efficient model for a specific prompt rather than using a single high-cost model for all tasks. Coinbase implemented this strategy to cut expenses nearly in half while significantly increasing their cache hit rate. Plano is an open-source tool that provides this routing layer, allowing developers to define routing preferences in plain English via YAML configurations. The system includes features like model affinity to maintain session consistency and cost-aware selection to prioritize cheaper models based on live pricing. Stanford researchers have introduced Shepherd, a runtime layer that provides Git-like version control for live AI agent environments. By snapshotting the agent's process, filesystem, and KV cache using copy-on-write forks, Shepherd allows agents to revert to previous states without losing progress or re-running expensive tool calls. This approach addresses the limitations of standard message logs, which fail to capture the underlying system state like installed packages or open file handles. Benchmarks show that this capability significantly improves success rates on complex coding tasks by allowing for automated error correction and state recovery.
閱讀原文 ↗目錄
How to reduce LLM costs by 50-60% using model routing
Model routing reduces LLM costs by automatically selecting the most efficient model for a specific prompt rather than using a single high-cost model for all tasks. Coinbase implemented this strategy to cut expenses nearly in half while significantly increasing their cache hit rate. Plano is an open-source tool that provides this routing layer, allowing developers to define routing preferences in plain English via YAML configurations. The system includes features like model affinity to maintain session consistency and cost-aware selection to prioritize cheaper models based on live pricing.
- Coinbase reduced LLM spend by nearly 50% and increased cache hit rates from 5% to 60% using model routing.
- Plano is an open-source routing layer that sits between applications and model providers to classify and forward prompts.
- Model affinity pins a specific model to a session ID for 10 minutes to prevent cache invalidation and context loss during multi-step tasks.
- Plano supports three routing methods: model-based, alias-based, and preference-aligned routing.
- Preference-aligned routing uses AI to infer the domain and action of a prompt to match it with the best model automatically.
- Cost-aware selection allows the router to pick the cheapest model among candidates that can handle a specific task equally well.
- Integrating agents like Hermes with Plano involves pointing the agent's OpenAI-compatible endpoint to the local Plano proxy.
Stanford researchers built the agent-native version of Git
Stanford researchers have introduced Shepherd, a runtime layer that provides Git-like version control for live AI agent environments. By snapshotting the agent's process, filesystem, and KV cache using copy-on-write forks, Shepherd allows agents to revert to previous states without losing progress or re-running expensive tool calls. This approach addresses the limitations of standard message logs, which fail to capture the underlying system state like installed packages or open file handles. Benchmarks show that this capability significantly improves success rates on complex coding tasks by allowing for automated error correction and state recovery.
- Shepherd snapshots the entire agent environment, including the process and filesystem, rather than just logging messages.
- The system utilizes copy-on-write forks that are five times faster than Docker commits for state management.
- Reverting states with Shepherd enables over 95% KV cache reuse, reducing token consumption and latency.
- In CooperBench testing, Shepherd increased the pass rate of a pair-coding agent setup from 28.8% to 54.7%.
- Unlike the Claude Code rewind feature, Shepherd can undo environment-level changes like pip installations and bash-driven file edits.
- External side effects such as emails or database writes require manual inverse operations as they cannot be automatically rolled back.