← 回到 Reading
Daily Dose of DS 2026-07-17

RLHF vs. DPO vs. GRPO in RL

Moonshot AI has released Kimi K3, a new model that outperforms Claude Opus 4.8 on several key benchmarks and competes with Claude Fable 5 and GPT-5.6 Sol. Kimi K3 is notable for its lower API pricing and its upcoming open-weights release, which will feature 2.8 trillion parameters. The release highlights the advantages of open-weights models over API-based "rented intelligence," which can be subject to regulatory shutdowns or pricing changes. This section compares three primary methods for aligning AI models: RLHF, DPO, and GRPO. RLHF is the traditional approach using four concurrent models and PPO, which is computationally expensive and difficult to stabilize. DPO simplifies the process by deriving rewards implicitly from preference pairs, while GRPO, introduced by DeepSeek, maintains the RL framework but replaces the critic with group-based statistics to reduce overhead. Paged Attention is a memory management technique for Large Language Models that addresses the inefficiencies of the KV cache. Traditional serving methods waste significant GPU memory through contiguous pre-allocation and redundant storage of shared prompts. By applying operating system concepts like virtual memory paging, Paged Attention allows for non-contiguous memory allocation and efficient sharing of data across requests. This optimization leads to significantly higher throughput and reduced memory fragmentation in production inference environments.

閱讀原文 ↗
目錄 3 段
  1. 01Kimi K3 vs Opus 4.8 vs Fable 5 vs GPT-5.6 Sol
  2. 02RLHF vs. DPO vs. GRPO in RL
  3. 03Paged Attention in LLMs
LLMs

Kimi K3 vs Opus 4.8 vs Fable 5 vs GPT-5.6 Sol

Moonshot AI has released Kimi K3, a new model that outperforms Claude Opus 4.8 on several key benchmarks and competes with Claude Fable 5 and GPT-5.6 Sol. Kimi K3 is notable for its lower API pricing and its upcoming open-weights release, which will feature 2.8 trillion parameters. The release highlights the advantages of open-weights models over API-based "rented intelligence," which can be subject to regulatory shutdowns or pricing changes.

  • Kimi K3 outperforms Claude Opus 4.8 on four out of five major benchmarks, including GPQA Diamond and FrontierSWE.
  • Kimi K3 is priced at approximately one-third the cost of Claude Fable 5's API.
  • Moonshot AI plans to release the 2.8 trillion parameters of Kimi K3 as open weights on July 27.
  • Claude Fable 5 was recently offline for 18 days due to a US export directive, illustrating the risks of API-dependent models.
  • Kimi K3 is expected to be the largest open-weights model released to date upon its full release.
RL

RLHF vs. DPO vs. GRPO in RL

This section compares three primary methods for aligning AI models: RLHF, DPO, and GRPO. RLHF is the traditional approach using four concurrent models and PPO, which is computationally expensive and difficult to stabilize. DPO simplifies the process by deriving rewards implicitly from preference pairs, while GRPO, introduced by DeepSeek, maintains the RL framework but replaces the critic with group-based statistics to reduce overhead.

  • RLHF requires four live models (policy, reference, reward, and critic), leading to high computational costs and stability challenges.
  • DPO (Direct Preference Optimization) removes the need for an explicit reward model and critic by using log-probability ratios from preference pairs.
  • GRPO (Group Relative Policy Optimization) was introduced by DeepSeek in 2024 to eliminate the critic bottleneck in RL training.
  • GRPO calculates advantages using group statistics (mean and standard deviation) rather than a value baseline from a critic model.
  • DPO can be brittle if the provided preference data does not sufficiently cover specific failure modes.
  • RLHF uses a KL penalty to ensure the trained policy remains close to the original reference model.
LLMs

Paged Attention in LLMs

Paged Attention is a memory management technique for Large Language Models that addresses the inefficiencies of the KV cache. Traditional serving methods waste significant GPU memory through contiguous pre-allocation and redundant storage of shared prompts. By applying operating system concepts like virtual memory paging, Paged Attention allows for non-contiguous memory allocation and efficient sharing of data across requests. This optimization leads to significantly higher throughput and reduced memory fragmentation in production inference environments.

  • Memory capacity is typically the primary bottleneck in scaling LLM inference, rather than compute power.
  • Traditional KV cache implementations often utilize only 20-30% of allocated GPU memory effectively due to fragmentation.
  • Paged Attention divides the KV cache into small, fixed-size blocks (typically 16 tokens) that can be stored non-contiguously.
  • A block table maps logical token indices to physical memory locations, decoupling the model's view from physical storage.
  • Shared prefixes, such as system prompts, can be stored once in physical memory and referenced by multiple requests via block tables.
  • The vLLM engine, which uses Paged Attention, can achieve 2-4x higher throughput compared to previous state-of-the-art systems.
  • Major inference frameworks including TensorRT-LLM and SGLang have adopted paging mechanisms similar to Paged Attention.