How LLM Routing Actually Works in Production
Transformer Lab is an open-source machine learning platform designed to orchestrate GPU resources across various cloud environments. It supports a wide range of training and evaluation workflows, including LoRA, DPO, and diffusion models, through a consistent interface. The platform is compatible with diverse hardware architectures, from Apple Silicon to NVIDIA and AMD clusters, and integrates with tools like Slurm and vLLM. LLM routing in production requires accounting for prompt caching to achieve actual cost savings. Since caches are model-specific, switching models mid-task resets the cache and incurs higher costs for the accumulated context. Production routers solve this using model affinity, which pins a model to a session ID for the duration of a task. The open-source tool Plano implements this via a four-stage pipeline including guardrails, the Arch-Router model, and selection policies. Prompt caching is an architectural strategy for AI agents that stores the mathematical state of static prompt prefixes to avoid redundant computation. By separating a prompt into a static prefix containing instructions and tools and a dynamic suffix for conversation history, systems can skip the expensive prefill phase for repeated content. This approach can reduce token costs by up to 90% for cached reads, as demonstrated by high-efficiency implementations like Claude Code. However, maintaining a high cache hit rate requires strict prompt ordering and avoiding any mutations to the prefix, such as dynamic timestamps or tool reordering.
閱讀原文 ↗目錄
The operating system for AI research labs!
Transformer Lab is an open-source machine learning platform designed to orchestrate GPU resources across various cloud environments. It supports a wide range of training and evaluation workflows, including LoRA, DPO, and diffusion models, through a consistent interface. The platform is compatible with diverse hardware architectures, from Apple Silicon to NVIDIA and AMD clusters, and integrates with tools like Slurm and vLLM.
- Transformer Lab supports multiple fine-tuning and optimization techniques including LoRA, QLoRA, DPO, ORPO, and SIMPO.
- The platform provides one-click conversion between Hugging Face, GGUF, and MLX model formats.
- It integrates with popular inference and training backends such as MLX, vLLM, Ollama, and Hugging Face Transformers.
- Evaluation capabilities include LLM-as-a-judge and the EleutherAI evaluation harness.
- The tool can submit jobs to cluster management systems like Slurm and SkyPilot from a single UI.
- It is compatible with diverse hardware architectures including Apple Silicon, NVIDIA, and AMD.
How LLM routing actually works in production
LLM routing in production requires accounting for prompt caching to achieve actual cost savings. Since caches are model-specific, switching models mid-task resets the cache and incurs higher costs for the accumulated context. Production routers solve this using model affinity, which pins a model to a session ID for the duration of a task. The open-source tool Plano implements this via a four-stage pipeline including guardrails, the Arch-Router model, and selection policies.
- Prompt caching reduces input token costs by approximately 90% for previously seen context.
- Prompt caches are model-specific, meaning switching models mid-task results in cold billing rates for the entire context.
- Model affinity pins a specific model to a session ID to keep the prompt cache warm across multiple calls in a task.
- A production routing layer typically includes four stages: Guardrail filter, Router model, Selection policy, and Model affinity.
- Arch-Router is a 1.5B parameter model trained on human preference data used to categorize and route prompts.
- Plano is an open-source proxy that implements this routing architecture to reduce costs without modifying agent code.
What is prompt caching?
Prompt caching is an architectural strategy for AI agents that stores the mathematical state of static prompt prefixes to avoid redundant computation. By separating a prompt into a static prefix containing instructions and tools and a dynamic suffix for conversation history, systems can skip the expensive prefill phase for repeated content. This approach can reduce token costs by up to 90% for cached reads, as demonstrated by high-efficiency implementations like Claude Code. However, maintaining a high cache hit rate requires strict prompt ordering and avoiding any mutations to the prefix, such as dynamic timestamps or tool reordering.
- Prompt caching stores Key and Value (KV) tensors computed during the prefill phase to skip redundant matrix multiplications in subsequent requests.
- LLM inference consists of a compute-bound prefill phase and a memory-bound decode phase; caching specifically optimizes the prefill phase.
- Cache hits are determined by a cryptographic hash of the token sequence, meaning any change in order or content results in a cache miss.
- Anthropic's pricing model offers a 90% discount for cache reads (0.1x base price) while charging a 25% premium for cache writes.
- Effective prompt architecture places static elements like system instructions and tool definitions at the top, with dynamic conversation history at the bottom.
- Claude Code serves as a primary example of high-efficiency caching, achieving over 90% hit rates and an 81% cost reduction in coding sessions.
- Cache efficiency can be monitored via specific API response fields: cache_creation_input_tokens and cache_read_input_tokens.