← 回到 Reading
Daily Dose of DS 2026-08-14

How Production LLMs Reason Better At Inference Time

Benchmarks from MCPMark V2 indicate that advanced Claude models can consume 54% more tokens than expected when interacting with traditional backends. This inefficiency stems from the model's attempt to resolve missing or unstructured backend context through excessive discovery queries and retries. While platforms like Supabase provide verbose documentation and ambiguous errors, the agent-centric backend InsForge uses structured JSON topology to reduce token load. Testing shows InsForge completed a RAG application build with significantly fewer tokens and zero manual interventions compared to Supabase. Inference-time compute scaling for LLMs involves parallel sampling or sequential trajectory extension to improve reasoning without altering model weights. While techniques like Majority Voting and Best-of-N leverage multiple outputs, sequential methods like Thinking Budgets and Self-Correction focus on internal processing. Advanced tree-based search methods like MCTS offer high performance but face challenges with reward model reliability and definition of step correctness. The focus of AI agent performance is shifting from the underlying LLM to the 'agent harness,' which manages orchestration, tools, and context. Benchmarks like SWE-bench Pro demonstrate that changes in the scaffold can have a 22x larger impact on performance than switching between frontier models. Building an effective harness involves seven key design decisions, including agent count, reasoning strategy, and verification methods. While models are becoming more capable and internalizing some planning logic, the harness remains a critical component for reliability and tool execution.

閱讀原文 ↗
目錄 3 段
  1. 01A smarter Claude model burns more tokens, not fewer
  2. 02How production LLMs reason better at inference time
  3. 037 decisions that define every Agent Harness
OPEN-SOURCE

A smarter Claude model burns more tokens, not fewer

Benchmarks from MCPMark V2 indicate that advanced Claude models can consume 54% more tokens than expected when interacting with traditional backends. This inefficiency stems from the model's attempt to resolve missing or unstructured backend context through excessive discovery queries and retries. While platforms like Supabase provide verbose documentation and ambiguous errors, the agent-centric backend InsForge uses structured JSON topology to reduce token load. Testing shows InsForge completed a RAG application build with significantly fewer tokens and zero manual interventions compared to Supabase.

  • Smarter models like Claude may increase token usage by 54% when forced to perform extensive discovery on unstructured backends.
  • Traditional backends often return 5-10x more tokens than necessary by providing full documentation instead of specific state data.
  • InsForge provides a full backend topology in approximately 500 tokens, significantly reducing the discovery phase.
  • Semantic exit codes and structured JSON in InsForge prevent the agent from entering infinite retry loops caused by ambiguous error codes.
  • In a comparative test, InsForge consumed 3.7M tokens versus Supabase's 10.4M tokens for the same RAG application build.
  • InsForge utilizes narrowly scoped skills to ensure the agent only loads context relevant to the current task.
LLMs

How production LLMs reason better at inference time

Inference-time compute scaling for LLMs involves parallel sampling or sequential trajectory extension to improve reasoning without altering model weights. While techniques like Majority Voting and Best-of-N leverage multiple outputs, sequential methods like Thinking Budgets and Self-Correction focus on internal processing. Advanced tree-based search methods like MCTS offer high performance but face challenges with reward model reliability and definition of step correctness.

  • Inference-time scaling is categorized into parallel scaling (multiple samples) and sequential scaling (longer trajectories).
  • Test-time compute can enable a smaller model to outperform a model 14x its size on specific tasks.
  • Majority voting relies on answer agreement and does not require a separate reward model, but it cannot fix consistent errors.
  • Best-of-N sampling is limited by the quality of the reward model, which can be exploited if optimization pressure is too high.
  • Self-correction loops can sometimes decrease accuracy, as seen with GPT-3.5 on GSM8K where it changed more correct answers than wrong ones.
  • DeepSeek R1 moved away from MCTS and Process Reward Models (PRMs) in favor of rule-based reinforcement learning (GRPO).
AGENTS

7 decisions that define every Agent Harness

The focus of AI agent performance is shifting from the underlying LLM to the 'agent harness,' which manages orchestration, tools, and context. Benchmarks like SWE-bench Pro demonstrate that changes in the scaffold can have a 22x larger impact on performance than switching between frontier models. Building an effective harness involves seven key design decisions, including agent count, reasoning strategy, and verification methods. While models are becoming more capable and internalizing some planning logic, the harness remains a critical component for reliability and tool execution.

  • Scaffold changes on SWE-bench Pro produce 22x larger performance swings than model swaps at the frontier.
  • Frontier models are saturating general benchmarks like MMLU, making infrastructure the primary differentiator.
  • Anthropic and OpenAI recommend a single-agent approach initially to avoid routing overhead and context loss.
  • Verification layers using computational checks like tests and linters provide deterministic ground truth for agents.
  • Tool scoping via lazy loading improves performance by only surfacing tools when they are relevant to the task.
  • Reasoning strategies like ReAct offer flexibility but are more costly than plan-then-execute approaches.
  • The industry trend shows harnesses getting thinner as models improve, though they remain necessary for context and tool management.