← 回到 Reading
Daily Dose of DS 2026-08-03

The Hands-on AI Engineer Playbook to Build RAG Apps for Production

The RAG Systems course has released a deep dive into the prefill stage of RAG applications, which is identified as the primary source of latency and cost. While retrieval and vector search are fast, processing retrieved tokens scales quadratically and can take several seconds. The guide explores techniques like KV cache reuse and selective recomputation to reduce Time to First Token (TTFT) and GPU costs. Tools like LMCache and TurboRAG are highlighted for their ability to significantly improve performance without sacrificing answer quality. The text explores the differences between Outcome Reward Models (ORM) and Process Reward Models (PRM) in verifying LLM reasoning. While PRMs are more effective at ranking solutions by evaluating individual steps, they are susceptible to model-based errors when used as training objectives. Consequently, rule-based outcome scoring is often preferred for training loops to ensure robustness, a strategy notably employed by DeepSeek-R1. This section outlines 16 specific techniques to optimize neural network training, focusing on improving speed and reducing memory consumption. It covers a range of strategies, from basic hardware acceleration and optimizer selection to advanced methods like Bayesian optimization for hyperparameter tuning. The text also provides practical implementation advice for distributed training and efficient data loading using tools like DeepSpeed and PyTorch's DistributedDataParallel.

閱讀原文 ↗
目錄 3 段
  1. 01The hands-on AI engineer playbook to build RAG apps for production
  2. 02ORM vs PRM: How LLMs verify their reasoning?
  3. 0316 techniques to optimize neural network training
AI ENGINEERING

The hands-on AI engineer playbook to build RAG apps for production

The RAG Systems course has released a deep dive into the prefill stage of RAG applications, which is identified as the primary source of latency and cost. While retrieval and vector search are fast, processing retrieved tokens scales quadratically and can take several seconds. The guide explores techniques like KV cache reuse and selective recomputation to reduce Time to First Token (TTFT) and GPU costs. Tools like LMCache and TurboRAG are highlighted for their ability to significantly improve performance without sacrificing answer quality.

  • Prefill latency is the dominant performance bottleneck in RAG systems, often taking seconds compared to milliseconds for retrieval.
  • The prefill stage scales quadratically with input length because every token must attend to every other token.
  • Standard prefix caching typically results in a near-zero hit rate for RAG workloads.
  • CacheBlend reduces Time to First Token (TTFT) by 2-3x by recomputing only 10-15% of tokens.
  • TurboRAG can achieve up to a 9.4x reduction in TTFT by moving the prefill process offline.
  • LMCache is an open-source package that implements CacheBlend and integrates with the vLLM serving engine.
REINFORCEMENT LEARNING

ORM vs PRM: How LLMs verify their reasoning?

The text explores the differences between Outcome Reward Models (ORM) and Process Reward Models (PRM) in verifying LLM reasoning. While PRMs are more effective at ranking solutions by evaluating individual steps, they are susceptible to model-based errors when used as training objectives. Consequently, rule-based outcome scoring is often preferred for training loops to ensure robustness, a strategy notably employed by DeepSeek-R1.

  • Process Reward Models (PRM) score every step of a solution, while Outcome Reward Models (ORM) only score the final answer.
  • In OpenAI's testing, a process scorer achieved 78.2% accuracy on a math test set, significantly higher than the outcome scorer's 72.4%.
  • PRMs are better for ranking and search because they can identify correct answers reached through faulty reasoning.
  • Rule-based outcome scoring is more reliable for training loops because it avoids the 'model gaps' and misjudgments inherent in model-based process scorers.
  • DeepSeek-R1 utilized rule-based outcome scoring to check final answers and output formats during its development.
  • Reinforcement learning techniques like PPO and GRPO are central to modern LLM post-training and reward modeling.
DEEP LEARNING

16 techniques to optimize neural network training

This section outlines 16 specific techniques to optimize neural network training, focusing on improving speed and reducing memory consumption. It covers a range of strategies, from basic hardware acceleration and optimizer selection to advanced methods like Bayesian optimization for hyperparameter tuning. The text also provides practical implementation advice for distributed training and efficient data loading using tools like DeepSpeed and PyTorch's DistributedDataParallel.

  • Bayesian optimization is more efficient than standard hyperparameter searches, achieving better F1 scores in fewer iterations.
  • Mixed precision training utilizes float16 for specific operations to speed up training while maintaining float32 for precision.
  • Activation checkpointing can reduce memory usage by a factor of sqrt(M) by recomputing intermediate activations during the backward pass.
  • DistributedDataParallel is recommended over DataParallel for all training scenarios to improve performance.
  • Normalizing data on the GPU instead of the CPU reduces memory bandwidth requirements by transferring smaller 8-bit integer types.
  • Creating tensors directly on the GPU using the device argument is significantly faster than transferring them from the CPU.
  • Configuring max_workers and pin_memory in DataLoaders allows for overlapping CPU data preparation with GPU computation.