Researchers Found a Way to Make LLMs 8.5x Faster!
Doc Holiday is an automation tool designed to eliminate documentation lag by integrating directly into the CI/CD pipeline. When a pull request is merged, the tool analyzes commit history, tickets, and specifications to automatically generate changelogs and release notes. This ensures that documentation remains synchronized with the codebase without requiring manual effort from development teams. The generated updates are automatically pushed to common collaboration platforms like Confluence, Notion, and Slack. DFlash is a new speculative decoding technique designed to overcome the speed limitations of traditional autoregressive draft models. By replacing the sequential drafter with a lightweight block diffusion model, DFlash can generate multiple tokens in parallel, maintaining a flat drafting cost regardless of the number of tokens speculated. The system further improves accuracy by conditioning the drafter on hidden features from the target model's layers. Benchmarks show DFlash achieving up to 415 tokens per second, an 8.5x improvement over vanilla decoding with no loss in output quality. Knowledge distillation is a technique used to transfer reasoning and knowledge from a larger 'Teacher' LLM to a smaller 'Student' LLM. This process can occur during pre-training, post-training, or both, allowing models to learn from each other rather than relying solely on raw text. The text highlights three primary distillation methods: matching full softmax probabilities, using one-hot output tokens, and joint training of both models simultaneously.
閱讀原文 ↗目錄
Documentation that keeps pace with the codebase
Doc Holiday is an automation tool designed to eliminate documentation lag by integrating directly into the CI/CD pipeline. When a pull request is merged, the tool analyzes commit history, tickets, and specifications to automatically generate changelogs and release notes. This ensures that documentation remains synchronized with the codebase without requiring manual effort from development teams. The generated updates are automatically pushed to common collaboration platforms like Confluence, Notion, and Slack.
- Doc Holiday automates the creation of changelogs, release notes, and documentation updates to prevent post-ship delays.
- The tool operates within the CI/CD pipeline and is triggered by pull request merges.
- It extracts information from commit history, linked tickets, and connected specifications to generate content.
- Outputs are compatible with major productivity tools including Confluence, Notion, Google Docs, and Slack.
- The system transforms documentation into a byproduct of the shipping process rather than a separate manual task.
Researchers found a way to make LLMs 8.5x faster!
DFlash is a new speculative decoding technique designed to overcome the speed limitations of traditional autoregressive draft models. By replacing the sequential drafter with a lightweight block diffusion model, DFlash can generate multiple tokens in parallel, maintaining a flat drafting cost regardless of the number of tokens speculated. The system further improves accuracy by conditioning the drafter on hidden features from the target model's layers. Benchmarks show DFlash achieving up to 415 tokens per second, an 8.5x improvement over vanilla decoding with no loss in output quality.
- Traditional speculative decoding is often bottlenecked by autoregressive draft models, limiting speedups to 2-3x.
- DFlash uses a block diffusion model to predict multiple tokens simultaneously in a single parallel shot.
- The drafting cost in DFlash remains constant even as the number of speculated tokens increases.
- DFlash incorporates hidden features from the target model into the draft layers to improve prediction accuracy.
- Experimental results demonstrate an 8.5x speed increase, moving from 48.5 to 415 tokens per second.
- DFlash is compatible with major frameworks including vLLM, SGLang, and Hugging Face Transformers.
Train LLMs using other LLMs
Knowledge distillation is a technique used to transfer reasoning and knowledge from a larger 'Teacher' LLM to a smaller 'Student' LLM. This process can occur during pre-training, post-training, or both, allowing models to learn from each other rather than relying solely on raw text. The text highlights three primary distillation methods: matching full softmax probabilities, using one-hot output tokens, and joint training of both models simultaneously.
- Distillation allows smaller models to achieve higher performance by learning from the outputs of larger, proprietary models.
- Llama 4 Scout and Maverick were trained using joint training with Llama 4 Behemoth during the pre-training stage.
- DeepSeek-R1 was distilled into Qwen and Llama 3.1 models using a post-training approach.
- Gemma 2 and 3 utilized Google's proprietary Gemini models for their training process.
- Matching full softmax probabilities is memory-intensive, potentially requiring 500 million GBs for a 5 trillion token corpus.
- Joint training involves training a teacher on hard labels while the student matches the teacher's softmax probabilities in the same batch.
Our AI Engineering, MCP, Agents, and DS books
The newsletter highlights four free books covering various aspects of AI and data science. These include an AI Engineering book focused on system design patterns, an MCP book with hands-on projects, an AI Agents book detailing agentic design patterns, and a comprehensive Data Science book. Each resource aims to provide practical knowledge and fundamentals for AI engineers and data scientists.
- Four free books are available covering AI Engineering, MCP, AI Agents, and Data Science.
- The AI Engineering book addresses system design patterns including prompt engineering, fine-tuning, and observability.
- The MCP book includes 11 hands-on projects focused on protocol fundamentals.
- The AI Agents book features 12 hands-on projects and covers agentic design patterns.
- The Data Science book is a 530-page guide covering classical ML, deep learning, and memory optimization.
- All books are hosted on Google Drive for free access.