A Cheaper Model Does Not Imply a Cheaper Turn
Alook is an open-source, self-hosted platform designed for multi-agent orchestration using a traditional organizational chart structure. Instead of manually wiring complex graphs, users define roles and reporting lines for agents who then communicate via email inboxes. This approach allows for intuitive local management of AI teams using existing agent runtimes like Claude Code and OpenCode. Model routing is a strategy used to reduce costs by directing simple tasks to cheaper LLMs, but it can be counterproductive in long agent sessions. Because prompt caches are model-specific, switching models mid-conversation forces a full 'cold prefill' of the entire history, losing the 90% discount provided by warm caches. In scenarios with large conversation histories, the cost of re-processing the history on a cheaper model often exceeds the savings gained from its lower per-token rate. The quality of Claude's output is primarily determined by prompt structure, which can be broken down into eight essential building blocks. These blocks include defining a role, specifying objectives with success criteria, and using XML tags to clearly delineate context and examples. Research from Anthropic suggests that placing long documents at the start of a prompt can boost performance by 30%. Advanced techniques like pre-filling and reasoning tags further refine the model's accuracy and formatting.
閱讀原文 ↗目錄
How to build your own AI company (100% local)
Alook is an open-source, self-hosted platform designed for multi-agent orchestration using a traditional organizational chart structure. Instead of manually wiring complex graphs, users define roles and reporting lines for agents who then communicate via email inboxes. This approach allows for intuitive local management of AI teams using existing agent runtimes like Claude Code and OpenCode.
- Alook replaces manual graph-based agent orchestration with a hierarchical organization chart structure.
- Agents in the Alook system coordinate and exchange information through dedicated email inboxes.
- The platform is designed to run 100% locally and is open-source, ensuring data privacy.
- Alook supports multiple agent runtimes including Claude Code, OpenCode, and Codex.
- The organizational structure allows for a single human point of contact with a 'CEO' agent who delegates tasks down the chain.
A cheaper model does not imply a cheaper turn
Model routing is a strategy used to reduce costs by directing simple tasks to cheaper LLMs, but it can be counterproductive in long agent sessions. Because prompt caches are model-specific, switching models mid-conversation forces a full 'cold prefill' of the entire history, losing the 90% discount provided by warm caches. In scenarios with large conversation histories, the cost of re-processing the history on a cheaper model often exceeds the savings gained from its lower per-token rate.
- Prompt caching allows providers to bill tokens at 10% of the base input rate when the prefix matches a stored KV cache.
- KV caches are model-specific and cannot be transferred between models like Opus 5 and Haiku 4.5 due to different weights.
- Switching models mid-session invalidates the warm cache, requiring a full re-prefill of the entire conversation transcript.
- For Anthropic's Opus 5 and Haiku 4.5, switching to the cheaper model only saves money if the history is less than 8 times the size of the new input.
- Model routing is most effective for independent prompts or short sessions where the history penalty is minimal.
- The most efficient time to switch models is during a context reset or compaction when the cache is already being invalidated.
The anatomy of a Claude prompt
The quality of Claude's output is primarily determined by prompt structure, which can be broken down into eight essential building blocks. These blocks include defining a role, specifying objectives with success criteria, and using XML tags to clearly delineate context and examples. Research from Anthropic suggests that placing long documents at the start of a prompt can boost performance by 30%. Advanced techniques like pre-filling and reasoning tags further refine the model's accuracy and formatting.
- A well-structured Claude prompt uses eight building blocks: Role, Objective, Context, Examples, Thinking, Constraints, Output Format, and Pre-filling.
- Defining a specific role in the system prompt significantly alters Claude's reasoning and communication style.
- Including success criteria ('so that') provides Claude with a benchmark to evaluate its own performance.
- Placing long documents at the top of the prompt and the query at the end can improve response quality by 30%.
- XML tags like <context>, <examples>, and <thinking> are essential for organizing complex information and separating reasoning from final answers.
- Pre-filling the start of a response via the API allows users to bypass preambles and enforce specific output formats.
What do AI companies think about Agent Harness?
The agent harness represents the infrastructure surrounding an AI model that enables it to function as an agent, and companies currently disagree on how complex this layer should be. Anthropic advocates for a thin harness that relies on model intelligence, whereas LangGraph uses a thick harness with explicit graph-based logic. While this infrastructure often serves as temporary scaffolding that is removed as models improve, models can become dependent on specific harness structures during training. Evidence from TerminalBench 2.0 suggests that optimizing the harness alone can significantly improve agent performance without changing the underlying model.
- The agent harness is the infrastructure that wraps a model to manage turns, tools, and memory.
- Anthropic's dumb loop approach assumes that as models get smarter, the need for complex infrastructure decreases.
- OpenAI's Agents SDK prioritizes a code-first approach using native Python instead of domain-specific languages.
- LangGraph uses a thick harness where every decision point and transition is explicitly defined as a node or edge in a graph.
- The scaffolding metaphor describes harness logic that is removed once a model is capable of handling those tasks internally.
- LangChain improved its ranking on TerminalBench 2.0 from outside the top 30 to rank 5 by only changing its infrastructure.
- Models trained with a specific harness in the loop may see performance degradation if that harness is modified.