← 回到 Reading
Daily Dose of DS 2026-08-09

A 10-week Roadmap to Run LLMs in Production

The ReAct agent pattern suffers from context accumulation where failed steps persist in the prompt, potentially distracting the model from its objective. In contrast, the Plan-and-Act pattern separates the agent's logic into a planner and an executor, allowing for better context management by stripping unnecessary data like raw HTML. Research on web navigation tasks shows that while a poor planner can degrade performance, a well-trained planner with dynamic replanning significantly outperforms standard ReAct. This approach requires more API calls but prevents the accumulation of errors in the agent's reasoning trace. This section outlines a comprehensive 10-week curriculum designed to help AI engineers master the production deployment and optimization of Large Language Models (LLMs). The roadmap covers fundamental concepts like the roofline model and paged attention, alongside practical experience with serving frameworks like vLLM and SGLang. The curriculum emphasizes observability, performance benchmarking, and advanced techniques such as continuous batching and quantization to build robust inference services. Machine learning systems should evolve through a structured four-phase process rather than starting with complex deep learning architectures. The process begins with non-ML heuristics to establish a baseline, followed by simple interpretable models to validate the end-to-end pipeline. Subsequent phases focus on refining these simple models through feature engineering and hyperparameter tuning before finally moving to complex architectures like transformers only when evidence justifies the added cost. This incremental approach reduces risk, improves debuggability, and aligns with MLOps best practices.

閱讀原文 ↗
目錄 3 段
  1. 01ReAct vs Plan-and-Act pattern in agents
  2. 02A 10-week roadmap to run LLMs in production
  3. 03Phases of ML Modeling
AGENTS

ReAct vs Plan-and-Act pattern in agents

The ReAct agent pattern suffers from context accumulation where failed steps persist in the prompt, potentially distracting the model from its objective. In contrast, the Plan-and-Act pattern separates the agent's logic into a planner and an executor, allowing for better context management by stripping unnecessary data like raw HTML. Research on web navigation tasks shows that while a poor planner can degrade performance, a well-trained planner with dynamic replanning significantly outperforms standard ReAct. This approach requires more API calls but prevents the accumulation of errors in the agent's reasoning trace.

  • ReAct agents keep all previous thoughts, actions, and observations in the prompt, leading to context bloat and attention competition.
  • Plan-and-Act divides tasks between a high-level planner and a grounded executor to manage context more efficiently.
  • On the WebArena-Lite benchmark, a ReAct-style executor scored 36.97%, whereas a properly trained planner with replanning reached 53.94%.
  • A poorly trained planner can reduce performance significantly, scoring only 20.60% in the same tests.
  • Replanning after every action is essential for recovering from failed steps that a static plan cannot handle.
  • The Plan-and-Act pattern increases operational costs by requiring one planner call for every executor step.
AI ENGINEERING

A 10-week roadmap to run LLMs in production

This section outlines a comprehensive 10-week curriculum designed to help AI engineers master the production deployment and optimization of Large Language Models (LLMs). The roadmap covers fundamental concepts like the roofline model and paged attention, alongside practical experience with serving frameworks like vLLM and SGLang. The curriculum emphasizes observability, performance benchmarking, and advanced techniques such as continuous batching and quantization to build robust inference services.

  • AI engineers should use the roofline model to understand why LLM decoding is typically memory-bound.
  • Mastering production serving requires deep dives into the schedulers and paged attention mechanisms of vLLM and SGLang.
  • Observability stacks using Grafana and Prometheus are essential for tracking metrics like TTFT, inter-token latency, and queue depth.
  • Optimization strategies include continuous batching, chunked prefill, speculative decoding, and quantization methods like AWQ and GPTQ.
  • The roadmap involves building an inference service that is load-tested with over 1000 concurrent requests and publicly benchmarked.
  • The curriculum is available as an open-source project on GitHub under the repository patchy631/time-to-first-token.
MLOps

Phases of ML Modeling

Machine learning systems should evolve through a structured four-phase process rather than starting with complex deep learning architectures. The process begins with non-ML heuristics to establish a baseline, followed by simple interpretable models to validate the end-to-end pipeline. Subsequent phases focus on refining these simple models through feature engineering and hyperparameter tuning before finally moving to complex architectures like transformers only when evidence justifies the added cost. This incremental approach reduces risk, improves debuggability, and aligns with MLOps best practices.

  • Unnecessary complexity in ML systems often leads to low utility and should be avoided until justified.
  • Phase 1 involves creating a non-ML baseline, such as a heuristic or rule, to set a minimum performance bar.
  • Phase 2 uses simple models like logistic regression or decision trees to validate data ingestion and serving pipelines.
  • Phase 3 focuses on maximizing the performance of existing simple models through feature engineering and hyperparameter tuning.
  • Phase 4 introduces complex models like deep neural networks only after simpler approaches are exhausted.
  • At every phase of development, the best model from the previous phase serves as the new baseline.
  • A well-tuned simple model or ensemble can often meet production requirements without the need for deep learning.