← 回到 Reading
ByteByteGo 2026-09-23

How to Customize a Model to Learn New Tricks

Model customization begins with prompting and few-shot examples to define constraints and output expectations. When required context or updated documentation is missing, retrieval-augmented generation (RAG) supplies external data directly into the request input without modifying model parameters. If recurring weaknesses persist or input token overhead becomes prohibitive, fine-tuning provides an alternative by updating model parameters directly through training on examples. While fine-tuning embeds default behaviors and internalizes facts, retrieval remains necessary when information changes frequently or requires traceable sourcing. Fine-tuning adapts an existing pretrained model to specific tasks using a targeted dataset, significantly reducing the training needed compared to pretraining from scratch. Supervised fine-tuning (SFT) uses input-target pairs to update model parameters via backpropagation and an optimizer based on measured loss. Reinforcement learning from human feedback (RLHF) addresses SFT's limitations by training a separate reward model based on human-ranked outputs. While full fine-tuning modifies every parameter and incurs heavy memory costs, efficient methods like LoRA and QLoRA allow adaptation without adjusting every parameter independently. LoRA (Low-Rank Adaptation) is a parameter-efficient fine-tuning method that minimizes the number of parameters updated during training. It keeps original model weights frozen and attaches small trainable adapter components to internal calculations. These adapters use pairs of compact low-rank matrices to compute additive adjustments to the model's intermediate representations. A configurable rank setting dictates adapter capacity and resource overhead, enabling targeted adaptation of attention and other layer transformations without altering the base model's foundation.

閱讀原文 ↗
目錄 7 段
  1. 01When Instructions and Information Are Not Enough
  2. 02How Fine-Tuning Changes an Existing Model
  3. 03LoRA: Learning a Smaller Set of Changes
  4. 04QLoRA: Reducing the Memory the Frozen Model Still Needs
  5. 05Turning the Techniques Into a Training Process
  6. 06Using the Specialization in an Application
  7. 07Conclusion

When Instructions and Information Are Not Enough

Model customization begins with prompting and few-shot examples to define constraints and output expectations. When required context or updated documentation is missing, retrieval-augmented generation (RAG) supplies external data directly into the request input without modifying model parameters. If recurring weaknesses persist or input token overhead becomes prohibitive, fine-tuning provides an alternative by updating model parameters directly through training on examples. While fine-tuning embeds default behaviors and internalizes facts, retrieval remains necessary when information changes frequently or requires traceable sourcing.

  • Prompting and few-shot prompting guide model behavior using input instructions and examples without altering model parameters.
  • Retrieval-augmented generation (RAG) incorporates external or updated information into a model's input at query time.
  • Relying on large numbers of prompt examples increases the amount of input that applications must process and maintain.
  • Fine-tuning trains a model on many examples to embed desired behaviors directly into learned parameters.
  • Retrieval remains preferred over fine-tuning when underlying information changes frequently or traceable answers are required.

How Fine-Tuning Changes an Existing Model

Fine-tuning adapts an existing pretrained model to specific tasks using a targeted dataset, significantly reducing the training needed compared to pretraining from scratch. Supervised fine-tuning (SFT) uses input-target pairs to update model parameters via backpropagation and an optimizer based on measured loss. Reinforcement learning from human feedback (RLHF) addresses SFT's limitations by training a separate reward model based on human-ranked outputs. While full fine-tuning modifies every parameter and incurs heavy memory costs, efficient methods like LoRA and QLoRA allow adaptation without adjusting every parameter independently.

  • Fine-tuning builds on capabilities acquired during pretraining, lowering the amount of training required for specialization.
  • Supervised fine-tuning (SFT) measures prediction loss against target responses and updates parameters using backpropagation and an optimizer.
  • Reinforcement learning from human feedback (RLHF) ranks multiple generated responses to train a separate reward model that scores quality, helpfulness, and safety.
  • Full fine-tuning allows every model parameter to be updated, requiring substantial GPU memory to store weights, update states, and intermediate activations.
  • LoRA and QLoRA provide efficient alternatives to full fine-tuning by avoiding independent adjustments to every parameter.

LoRA: Learning a Smaller Set of Changes

LoRA (Low-Rank Adaptation) is a parameter-efficient fine-tuning method that minimizes the number of parameters updated during training. It keeps original model weights frozen and attaches small trainable adapter components to internal calculations. These adapters use pairs of compact low-rank matrices to compute additive adjustments to the model's intermediate representations. A configurable rank setting dictates adapter capacity and resource overhead, enabling targeted adaptation of attention and other layer transformations without altering the base model's foundation.

  • LoRA freezes original model weights and updates only small trainable adapter components attached to internal computations.
  • Each adapter is implemented via two compact matrices that create an intermediate representation and produce additive adjustments.
  • The rank hyperparameter dictates adapter capacity, balancing adaptation flexibility against parameter size and compute costs.
  • Standard LoRA adapters begin training with zero contribution, ensuring the model retains its original behavior at initialization.
  • LoRA can target specific internal operations, such as attention calculations and layer transformations.

QLoRA: Reducing the Memory the Frozen Model Still Needs

While LoRA reduces parameter update costs, the frozen base model still consumes significant memory for large models. QLoRA solves this problem by combining LoRA with quantization, typically storing the frozen base weights in a 4-bit format while keeping trainable adapters in 16-bit or 32-bit precision. Compressed weights are dynamically reconstructed to higher precision during calculation so adapters can train on the combined output. This approach allows larger models to be fine-tuned within a fixed memory budget, although total memory is still affected by activations and workspace.

  • QLoRA combines LoRA with weight quantization to reduce the memory footprint of the frozen base model.
  • QLoRA commonly stores frozen base weights in 4-bit precision while maintaining adapter parameters in 16-bit or 32-bit representations.
  • Base weights are stored in a compact format and reconstructed into higher-precision approximations during computation.
  • Adapters learn alongside the quantized base model, which allows them to adapt to and partially recover performance lost to quantization error.
  • Compressed weights only account for part of the memory profile; optimizer states, intermediate activations, and working space still require capacity.

Turning the Techniques Into a Training Process

Setting up an effective fine-tuning process requires selecting a capable base model and establishing a prompt baseline before training begins. Training data must be rigorously prepared and partitioned into distinct training, validation, and test splits to prevent data leakage and contradictory signals. Key configuration parameters such as rank, adapter placement, learning rate, batch size, gradient accumulation, and gradient checkpointing manage capacity, memory constraints, and stability. Finally, evaluation must track task-specific correctness and ensure that adapter training does not degrade other essential capabilities.

  • A baseline should be established with a prompt before training to measure improvement accurately.
  • Data must be strictly split into training, validation, and test sets without duplicate or related examples leaking across splits.
  • Contradictory labels or unsupported summary claims in the training data directly degrade learned behavior.
  • Gradient accumulation and gradient checkpointing reduce memory requirements during training and complement LoRA or QLoRA.
  • Overfitting occurs when training performance continues to improve while validation performance worsens, often requiring an earlier checkpoint.
  • Fine-tuning adapters can degrade existing capabilities even while keeping original weights frozen, necessitating end-to-end evaluation.

Using the Specialization in an Application

LoRA training produces lightweight adapter weights that can remain separate from the base model or be merged directly into it. Keeping adapters separate enables a single base model to serve multiple tasks, whereas merging eliminates runtime adapter overhead at the expense of higher storage. Implementation details for merging and quantization can alter numerical representations and outputs, requiring evaluation of the exact deployment version. Specialized models can continue to operate alongside standard application patterns like prompts, RAG, and output validation.

  • LoRA adapter files are significantly smaller than full models but require a compatible base checkpoint to function.
  • A single base model can be combined with different adapters to efficiently serve multiple specialized tasks.
  • Merging adapter weights into the base model creates a standalone full model, eliminating separate adapter computations at the cost of storage.
  • Quantization and merging behavior vary across serving implementations and can alter numerical representations and model outputs.
  • Specialized fine-tuned models can be combined with prompts, RAG pipelines, and output validation code.

Conclusion

Fine-tuning enables models with broad capabilities to achieve reliable specialization by learning from demonstrations and recurring response patterns. Parameter-efficient fine-tuning approaches like LoRA reduce computational costs by training compact adjustments over frozen original weights. QLoRA builds upon this by storing those frozen base weights at lower precision to significantly decrease memory consumption. However, these techniques must be paired with representative training data, clear task definitions, and thorough evaluation to achieve measurable improvements.

  • Fine-tuning adapts broadly capable models toward reliable specialization via demonstration data.
  • LoRA reduces fine-tuning costs by training compact weight adjustments while leaving original model weights frozen.
  • QLoRA further cuts memory overhead by quantizing the frozen weights to lower precision.
  • Neither LoRA nor QLoRA substitutes for quality examples, well-defined tasks, or thorough evaluation.