How to Shrink a Language Model Without Making it Too Dumb
Large language models derive their capabilities primarily from massive collections of numerical weights rather than traditional programmatic rules. For instance, a 70-billion-parameter model requires around 140 GB of storage, with weights structured into matrices distributed across dozens of layers. While supporting components like the Transformer architecture and context window are essential, weights store the fundamental patterns, grammar, and reasoning. Intelligence emerges from the relationships across these distributed weights, and because many weights are near zero with minimal output impact, shrinking models primarily involves optimizing these weights. Techniques for reducing model size focus on utilizing fewer bits per weight or decreasing the total weight count. The three primary methods are quantization, pruning, and knowledge distillation, each approaching compression via precision reduction, weight elimination, or architecture mimicry. These approaches are complementary and can be stacked sequentially across development and deployment pipelines. Ultimately, combining these methods allows complex large language models to run efficiently on standard consumer hardware. Quantization shrinks model size by reducing the precision allocated to each individual weight. While training requires 32-bit floating-point precision (FP32) for fine adjustments, models are typically distributed in 16-bit formats such as FP16 or BF16. Floating-point numbers allocate bits across sign, exponent, and mantissa components, with BF16 preserving the FP32 dynamic range by retaining 8 exponent bits while reducing the mantissa. Further quantization converts floating-point values into integers, such as 8-bit or 4-bit representations, which removes both precision and built-in scaling.
閱讀原文 ↗目錄
- 01What Makes a Language Model Intelligent?
- 02Three Techniques to Shrink the Model
- 03Shrinking a Model by Packing Fewer Details
- 041 - Mapping Ranges
- 052 - Rounding Values
- 063 - Using a Scale Factor
- 07Shrinking the Model by Removing Unused Weights
- 08Shrinking the Model by Mimicking Behaviour
- 09Does Shrinking Damage the Model’s Intelligence?
- 10Conclusion
What Makes a Language Model Intelligent?
Large language models derive their capabilities primarily from massive collections of numerical weights rather than traditional programmatic rules. For instance, a 70-billion-parameter model requires around 140 GB of storage, with weights structured into matrices distributed across dozens of layers. While supporting components like the Transformer architecture and context window are essential, weights store the fundamental patterns, grammar, and reasoning. Intelligence emerges from the relationships across these distributed weights, and because many weights are near zero with minimal output impact, shrinking models primarily involves optimizing these weights.
- A 70-billion-parameter model requires approximately 140 GB of storage when parameters are stored in standard 16-bit (2-byte) format.
- Weights are organized into matrices (such as 4096 by 4096 grids) stacked across numerous layers in an architecture.
- The Transformer architecture routes data and performs attention, but the weights constitute the bulk of the model file and its learned intelligence.
- Individual weights hold no standalone semantic meaning; knowledge is distributed across the relationships among weights.
- Model compression relies on modifying weights because most weights are close to zero and exert little impact on final outputs.
Three Techniques to Shrink the Model
Techniques for reducing model size focus on utilizing fewer bits per weight or decreasing the total weight count. The three primary methods are quantization, pruning, and knowledge distillation, each approaching compression via precision reduction, weight elimination, or architecture mimicry. These approaches are complementary and can be stacked sequentially across development and deployment pipelines. Ultimately, combining these methods allows complex large language models to run efficiently on standard consumer hardware.
- Model shrinking revolves around using fewer bits to store a weight or using fewer weights overall.
- Quantization retains all weights but describes each in a less precise format, such as reducing two bytes to half a byte.
- Pruning identifies and deletes weights that contribute nothing to the model.
- Knowledge distillation trains a separate, smaller model to mimic the behavior of the untouched original model.
- Compression techniques can be stacked sequentially by different actors across creation, research, and deployment stages.
- Stacking compression techniques enables high-end large language models to run on normal consumer hardware.
Shrinking a Model by Packing Fewer Details
Quantization shrinks model size by reducing the precision allocated to each individual weight. While training requires 32-bit floating-point precision (FP32) for fine adjustments, models are typically distributed in 16-bit formats such as FP16 or BF16. Floating-point numbers allocate bits across sign, exponent, and mantissa components, with BF16 preserving the FP32 dynamic range by retaining 8 exponent bits while reducing the mantissa. Further quantization converts floating-point values into integers, such as 8-bit or 4-bit representations, which removes both precision and built-in scaling.
- Quantization reduces model storage by discarding unnecessary numerical precision from weights.
- Training typically requires 32-bit floating-point numbers (FP32) to handle millions of tiny weight adjustments.
- Distributed models frequently reduce precision to 16-bit floating-point formats such as FP16 or BF16.
- BF16 uses 1 sign bit, 8 exponent bits, and 7 mantissa bits, preserving the dynamic range of FP32 with less precision.
- FP16 allocates 5 bits to the exponent and more to the mantissa compared to BF16, but BF16 has largely replaced FP16 in practice.
- Converting floats to 8-bit or 4-bit integers removes both precision and inherent scale.
1 - Mapping Ranges
The first step in quantizing weights involves determining the minimum and maximum values across a localized block of neighboring weights. Rather than scaling an entire model at once, this mapping operates over small groups to preserve precision. By identifying the maximum absolute magnitude within the block, the total span is partitioned into a fixed number of steps defined by the target precision. For instance, a 4-bit precision target allows seven steps in positive and negative directions, yielding a concrete step size calculated directly from the peak value.
- Range mapping begins by identifying the minimum and maximum values of a small group of neighboring weights called a block.
- The data set used for step calculation is restricted to small blocks rather than the entire model.
- A target precision of 4 bits provides a whole number range from -7 to 7, corresponding to seven steps in each direction.
- Step size is calculated by dividing the maximum absolute weight by the number of directional steps (for example, 0.070 divided by 7 equals 0.010).
2 - Rounding Values
This section explains the process of discretizing weights by rounding them to the closest available step size. Weights are divided by the step size and rounded to the nearest integer, producing values restricted between -7 and 7 for the model file. Because floating-point numbers inherently carry their own scale, rounding them removes both precision and scale. As a result, the removed scale factor must be retained and stored separately.
- Weights are converted to integers by dividing by a step size and rounding to the nearest whole number.
- The resulting values in the example map to whole numbers between -7 and 7 before entering the final model weights file.
- Original weights are represented as floating-point numbers, which inherently encode their own scale.
- Rounding weights eliminates both their decimal precision and individual scale.
- The scale discarded during rounding must be stored separately to retain proper magnitude.
3 - Using a Scale Factor
Tracking a scale factor enables compressed integers to reconstruct approximate original values when a model reads them. In this approach, a single scale factor, such as a step size, is stored once per block. Weights are recovered by multiplying each stored integer by the scale factor. While a minor reconstruction error remains, the overall model layers and matrices remain intact with significantly reduced storage requirements.
- A scale factor is required to reconstruct original values from compressed integer representations.
- The scale factor is stored once per block, functioning as the quantization step size.
- Decompressing a weight involves multiplying the stored integer by the block's scale factor.
- Reconstruction introduces minor numerical error compared to original values while dramatically lowering storage needs.
Shrinking the Model by Removing Unused Weights
Model pruning reduces model size by deleting weights that have minimal impact on the final output. While simple magnitude-based pruning sorts and removes weights closest to zero, more sophisticated approaches evaluate weight importance by monitoring inputs across sample texts. Pruning mechanisms generally fall into two categories: zeroing out weights, which leaves computation overhead on GPUs, or removing entire structural components like neurons, attention heads, or layers, which physically shrinks matrices at the risk of higher performance degradation. Pruning is typically combined with other optimization techniques rather than used alone.
- Pruning deletes non-essential weights that have near-zero values and do not meaningfully affect output.
- Sorting weights by magnitude is the simplest pruning method, but scoring weights based on sample input pathways yields better results.
- Zeroing out weights maintains model quality better but yields sparse matrices that GPUs must still process.
- Structural pruning eliminates full neurons, attention heads, or layers, reducing matrix dimensions but causing coarser damage to model capability.
- Pruning is rarely sufficient on its own and works best when combined with complementary compression techniques.
Shrinking the Model by Mimicking Behaviour
Knowledge distillation is a compression technique where a smaller student model is trained to mimic the behavior of a larger teacher model. Instead of relying strictly on raw text or one-hot correct answers, the student model is trained on the entire probability distribution generated by the teacher across vocabulary tokens. This allows the student to learn richer contextual nuances, such as plausible alternative word choices. Although running the teacher model during distillation demands datacenter-level resources, the resulting student model can be deployed on substantially less demanding hardware.
- Knowledge distillation compresses models by training a newly initialized student model to mimic a large teacher model's outputs.
- A student model typically retains the architecture of the teacher but with fewer layers and smaller weight matrices.
- Instead of hard targets, the student model receives the complete probability distribution of vocabulary scores generated by the teacher.
- Training via distillation requires datacenter-scale compute to process large volumes of data through the teacher model.
- Developers benefit from distillation by obtaining smaller models capable of running on less demanding hardware.
Does Shrinking Damage the Model’s Intelligence?
Shrinking a model involves a trade-off between physical size and overall mental sharpness, though intelligence degradation is typically small. Quantization reduces weight precision, causing minimal loss when moving from 32-bit to 8-bit, but major degradation at 4-bit or lower. Pruning affects complex, multi-step logic primarily when deeper pathways are aggressively removed rather than just trimming idle pathways. Knowledge distillation allows a smaller student model to mimic a teacher model's style, but the student often lacks original problem-solving skills for novel logical puzzles.
- Model shrinking causes a trade-off between physical size and mental sharpness, though overall intelligence loss is usually minor.
- Quantizing from 32-bit to 8-bit causes almost no noticeable change in intelligence, whereas dropping to 4-bit or lower has a significant impact.
- Quantization can degrade a model's grasp of nuance, recall of specific facts, and tone naturalness.
- Trimming idle pathways via pruning does not strongly impact intelligence, but aggressively pruning deeper pathways degrades multi-step logical reasoning.
- Knowledge distillation produces student models that mimic the teacher's style well but struggle with unfamiliar logical problems.
Conclusion
This conclusion reviews key techniques for shrinking large language models without significantly compromising their capabilities. The primary methods outlined include quantization, which reduces the bit-width of weights; pruning, which removes the least important weights and pathways; and knowledge distillation, which trains a smaller student model using a larger teacher model. Selecting the right technique or combination depends on the specific operational and performance goals required for the model.
- Compressing a language model without a structured approach can cause a substantial loss of model intelligence.
- Quantization compresses models by storing weights using fewer bits.
- Pruning reduces model size by eliminating weights and pathways that contribute minimally.
- Distillation preserves the original model and uses it as a teacher to train a more compact student model.
- The choice of shrinking method or combination depends on the target requirements of the language model.