How Big Models Teach Small Models to Be Smart
Distillation trains a new, separate student model to replicate the outputs and behavior of a larger teacher model. Unlike compression techniques such as quantization and pruning, which reduce the size of an existing model, distillation produces an entirely independent model with distinct parameters. The resulting smaller model provides practical benefits, including lower inference costs, decreased latency, and the ability to run on-device. This approach is exemplified by Google's Gemma models, which are distilled from the larger Gemini model family. Learning from a model's output provides richer information than standard hard labels because output distributions, termed soft labels, assign probabilities across all candidate options. These distributions capture relationships between classes, a concept often referred to as dark knowledge. During distillation, a student model learns to match the teacher's full confidence profile rather than just targeting single correct answers, which allows it to reach high performance using fewer training examples. A temperature control is employed to spread out probability values and further expose this fine structure for the student model. Knowledge distillation generally takes three forms depending on what the student model replicates: output distillation, feature distillation, and synthetic data distillation. Output distillation matches the teacher's final output probabilities, while feature distillation aligns intermediate internal representations, as used by Google to train EmbeddingGemma from Gemini. Synthetic data distillation fine-tunes a student on teacher-generated data, an approach popularized by Stanford's Alpaca. Because modern closed-source models often restrict access to internal values and probabilities, synthetic data distillation has become the most widely adopted variant in practice.
閱讀原文 ↗Distillation
Distillation trains a new, separate student model to replicate the outputs and behavior of a larger teacher model. Unlike compression techniques such as quantization and pruning, which reduce the size of an existing model, distillation produces an entirely independent model with distinct parameters. The resulting smaller model provides practical benefits, including lower inference costs, decreased latency, and the ability to run on-device. This approach is exemplified by Google's Gemma models, which are distilled from the larger Gemini model family.
- Distillation trains a distinct student model to replicate a teacher model's behavior rather than modifying the original model directly.
- Compression methods like quantization and pruning shrink an existing model by lowering precision or removing parts, whereas distillation trains a fresh model.
- Distilled models offer lower latency, reduced cost per request, and the capability to operate on local hardware such as phones.
- Google's Gemma models are created using distillation from the Gemini family of models.
- Distillation and quantization can be chained sequentially to optimize a model for specific target devices.
Soft Labels
Learning from a model's output provides richer information than standard hard labels because output distributions, termed soft labels, assign probabilities across all candidate options. These distributions capture relationships between classes, a concept often referred to as dark knowledge. During distillation, a student model learns to match the teacher's full confidence profile rather than just targeting single correct answers, which allows it to reach high performance using fewer training examples. A temperature control is employed to spread out probability values and further expose this fine structure for the student model.
- Soft labels provide a complete probability distribution over categories rather than a single discrete label.
- Dark knowledge refers to the structural relationships and relative plausibility among classes captured within teacher model probabilities.
- Student models in distillation are trained by minimizing the divergence between their output distribution and the teacher's soft labels.
- Training on soft targets enables student models to achieve strong performance from significantly fewer examples.
- Increasing the temperature parameter spreads out probability distributions, exposing finer category structures to the student model.
Methods
Knowledge distillation generally takes three forms depending on what the student model replicates: output distillation, feature distillation, and synthetic data distillation. Output distillation matches the teacher's final output probabilities, while feature distillation aligns intermediate internal representations, as used by Google to train EmbeddingGemma from Gemini. Synthetic data distillation fine-tunes a student on teacher-generated data, an approach popularized by Stanford's Alpaca. Because modern closed-source models often restrict access to internal values and probabilities, synthetic data distillation has become the most widely adopted variant in practice.
- Distillation is split into three main variants: output distillation, feature distillation, and synthetic data distillation.
- Output distillation dates to 2015 and requires access to the teacher's soft labels and probability distributions.
- Feature distillation matches internal model representations; Google applied this method to train EmbeddingGemma using Gemini representations.
- Synthetic data distillation fine-tunes a model on generated datasets, as demonstrated early on by Stanford's Alpaca.
- Access constraints on closed models make synthetic data distillation the most common method in practice, as it requires only textual output.
- Distillation strategies can be combined in a single training run, such as pairing synthetic data with soft labels.
Results
In early 2025, DeepSeek demonstrated that fine-tuning smaller models on outputs from a large reasoning model yielded remarkable performance on specific benchmarks. Notably, a 7-billion-parameter distilled student outperformed a 32-billion-parameter model in competition mathematics while remaining small enough to run locally on a single GPU. However, an important qualifier is that these gains are largely confined to narrow, well-defined domains such as math and code. Across broader measures of general knowledge, smaller distilled models consistently trail their larger counterparts.
- DeepSeek fine-tuned smaller existing models on training examples generated by a large reasoning model in early 2025.
- A 7-billion-parameter distilled model outperformed a 32-billion-parameter model on a competition mathematics benchmark.
- The released distilled model family ranged from 1.5 billion to 70 billion parameters, enabling local execution on a single graphics card for smaller variants.
- Distillation enables strong task-specific performance on narrow tasks like math and code cheaply and locally.
- Smaller distilled models still lag behind larger models on broader measures of general knowledge.
Limits
Knowledge distillation has notable constraints that influence its suitability for specific machine learning problems. Student models are bounded by a ceiling effect where their performance rarely exceeds the teacher, inheriting errors alongside accurate outputs. Furthermore, an excessively wide capacity gap between teacher and student can degrade transfer quality, sometimes necessitating intermediate models. Model architecture often outweighs raw parameter count in determining student success, and subtle behavioral traits can inadvertently transfer from teacher to student even through filtered, task-unrelated data.
- A student model's quality is typically capped at or below the teacher's level on seen data distributions, directly inheriting teacher errors.
- A very wide capacity gap between a large teacher and a small student can degrade transfer performance.
- Multi-step distillation using an intermediate-sized model can mitigate transfer degradation caused by wide capacity gaps.
- Base model architecture can be more influential than parameter count, as demonstrated by a 32B student outperforming a 70B student.
- Distillation can unintentionally transfer latent behavioral preferences and traits between models sharing the same base architecture, even when training data consists only of filtered number sequences.
Automation
Recent advances in knowledge distillation focus on fully automating the process using a self-running loop driven by a large teacher model. In this setup, the teacher autonomously generates training data, fine-tunes the student, and evaluates performance iteratively until progress plateaus, reducing human involvement primarily to setup and final verification. Recent research in 2026 demonstrated that while this method cuts manual pipeline and dataset creation efforts, the initial selection of the teacher model plays a decisive role in the final student model's quality.
- Automated distillation loops allow a large teacher model to generate data, fine-tune the student, and evaluate progress autonomously.
- Human responsibilities are restricted to defining the task and success criteria at the start and conducting a final evaluation on real data.
- 2026 research on a detection task showed that the choice of teacher model has a major impact on student model performance under identical loop settings.
- The automated process removes the requirement to manually assemble large hand-labeled datasets prior to training small, task-specific models.
Conclusion
Distillation is a training technique where a smaller student model learns to replicate the behavior of a larger teacher model rather than directly compressing it. It leverages soft labels that provide richer confidence distributions across outputs, often utilizing teacher-generated datasets. The effectiveness of distillation is subject to clear constraints, including the teacher's performance ceiling, negative effects from large size discrepancies, and the unintended transfer of behavioral traits. Consequently, the approach is best suited for well-defined, narrow tasks rather than broad, general-purpose capabilities.
- Distillation creates a separate, smaller model rather than compressing the original architecture directly.
- Soft labels facilitate distillation by providing information about the teacher's confidence distribution across classes.
- In practice, distillation is most commonly executed by having the teacher model generate training sets for the student.
- A wider size disparity between teacher and student models can degrade performance rather than improve it.
- Distillation can inadvertently transfer unintended behavioral traits from teacher to student models.
- Distillation is better suited for narrow, well-defined tasks than for general, open-ended capabilities.