← 回到 Reading
Daily Dose of DS 2026-06-24

Speculation Is All You Need!

Master.dev and Anthropic have partnered to release a free educational course on Claude Code. Taught by Lydia Hallie, a member of the Claude Code team at Anthropic, the course aims to explain the tool's internal mechanics. The live version of the course previously set platform records with over 10,000 attendees. It is now available to the public without requiring a subscription or trial. Speculative decoding typically speeds up LLM inference by 2-3x, but is often limited by the autoregressive nature of the small drafter model. Modal's new DFlash draft models overcome this bottleneck by using block diffusion to generate token proposals in a single parallel pass. By incorporating hidden states from the target model and training on production traffic, DFlash increases the acceptance length of proposed tokens. This approach allowed Qwen 3.5 122B-A10B to reach speeds of over 1000 tokens per second on a B200 GPU. ART (Agent Reinforcement Trainer) is an open-source framework designed to simplify reinforcement learning for LLM agents by eliminating manual reward engineering. It utilizes an LLM judge to grade multiple task attempts relatively, leveraging the Group Relative Policy Optimization (GRPO) algorithm. Unlike standard RL frameworks, ART supports complex agentic workflows including tool calls and multi-turn reasoning. By integrating tools like vLLM and Unsloth, it enables local fine-tuning of open-source models to achieve high performance on specific tasks.

閱讀原文 ↗
目錄 3 段
  1. 01Free Claude Code course with Lydia Hallie, Anthropic & Master.dev
  2. 02Speculation is all you need!
  3. 03Build Agents that can learn like humans
COURSES

Free Claude Code course with Lydia Hallie, Anthropic & Master.dev

Master.dev and Anthropic have partnered to release a free educational course on Claude Code. Taught by Lydia Hallie, a member of the Claude Code team at Anthropic, the course aims to explain the tool's internal mechanics. The live version of the course previously set platform records with over 10,000 attendees. It is now available to the public without requiring a subscription or trial.

  • Master.dev and Anthropic partnered to make the Claude Code course free for all users.
  • The course is taught by Lydia Hallie, who works on the Claude Code team at Anthropic.
  • The live session of the course broke platform records with more than 10,000 participants.
  • The curriculum focuses on visualizing how the tool works under the hood to improve user direction of AI.
  • Access to the course does not require a subscription or trial period.
LLMs

Speculation is all you need!

Speculative decoding typically speeds up LLM inference by 2-3x, but is often limited by the autoregressive nature of the small drafter model. Modal's new DFlash draft models overcome this bottleneck by using block diffusion to generate token proposals in a single parallel pass. By incorporating hidden states from the target model and training on production traffic, DFlash increases the acceptance length of proposed tokens. This approach allowed Qwen 3.5 122B-A10B to reach speeds of over 1000 tokens per second on a B200 GPU.

  • Standard speculative decoding is bottlenecked by autoregressive drafters that generate one token at a time.
  • DFlash replaces autoregressive drafting with block diffusion, allowing for parallel token generation in a single pass.
  • DFlash improves proposal accuracy by feeding hidden states from the target model's layers into the draft model.
  • Acceptance length is the primary driver of speedup because LLM decoding is memory-bound rather than compute-bound.
  • Training drafters on specific target model outputs and production traffic can significantly increase the acceptance length from a baseline of 3 to over 9.
  • DFlash is compatible with popular inference frameworks including vLLM, SGLang, and Hugging Face Transformers.
HANDS-ON

Build Agents that can learn like humans

ART (Agent Reinforcement Trainer) is an open-source framework designed to simplify reinforcement learning for LLM agents by eliminating manual reward engineering. It utilizes an LLM judge to grade multiple task attempts relatively, leveraging the Group Relative Policy Optimization (GRPO) algorithm. Unlike standard RL frameworks, ART supports complex agentic workflows including tool calls and multi-turn reasoning. By integrating tools like vLLM and Unsloth, it enables local fine-tuning of open-source models to achieve high performance on specific tasks.

  • ART eliminates the need for manual reward functions by using an LLM judge to compare task attempts.
  • The framework is based on Group Relative Policy Optimization (GRPO), the same algorithm used for DeepSeek R1.
  • ART supports complex agent behaviors such as multi-turn conversations and API tool calls.
  • It integrates with popular open-source tools including vLLM for model serving and Unsloth for training.
  • The framework allows small open-source models like Qwen to be fine-tuned to outperform larger closed-source models on specific tasks.
  • ART provides native integrations with agent frameworks like LangGraph, CrewAI, and ADK.