← 回到 Reading
ByteByteGo 2026-09-21

How to Run a Big Model on Cheap Hardware?

An AI model relies on numerical parameters, also known as weights, that are structured into layers to convert inputs into outputs. In language models, input text is broken into tokens, and the model iteratively generates responses by predicting probabilities for subsequent tokens. While training is a computationally intensive process that updates weights on expensive infrastructure, inference utilizes fixed weights and can often be executed on standard consumer hardware. Fitting large language models on smaller hardware is constrained by total memory capacity, memory location, and processing bottlenecks. An 8-billion-parameter model at 16-bit precision requires about 16 GB just for weights, plus extra space for software overhead, temporary calculations, and context. Because persistent storage like SSDs is too slow for execution, weights must reside in fast working memory such as RAM or GPU VRAM. Beyond capacity, memory bandwidth and compute architecture pose limits, with GPUs providing superior parallelism over CPUs and distinct bottlenecks arising across the prefill and decoding stages of inference. Local execution allows developers to utilize existing hardware for experimentation and decrease reliance on rented cloud infrastructure. This approach offers enhanced privacy for sensitive documents and proprietary code, alongside offline capability for desktop assistants and remote deployments. Additionally, running models locally provides greater control over model versions and application behavior. However, it is not universally cost-effective, as hardware, electricity, maintenance, and latency must be weighed against usage frequency compared to hosted services.

閱讀原文 ↗
目錄 12 段
  1. 01What Running a Model Actually Means
  2. 02What Makes it Difficult to Fit the Model on Smaller Hardware
  3. 03Why Make the Effort to Run Locally?
  4. 04Quantization: Give Each Weight a Smaller Representation
  5. 05Layer-Wise Offloading: Move Weights as They Are Needed
  6. 06Mixture of Experts: Use Selected Parts for Each Token
  7. 07Distillation: Let a Larger Model Teach a Smaller One
  8. 08Pruning: Remove Work That Contributes Less
  9. 09The Conversation Has Its Own Memory Requirements
  10. 10Better Software Can Make the Same Hardware More Useful
  11. 11Speculative Decoding: Propose Several Tokens Before Checking Them
  12. 12Conclusion

What Running a Model Actually Means

An AI model relies on numerical parameters, also known as weights, that are structured into layers to convert inputs into outputs. In language models, input text is broken into tokens, and the model iteratively generates responses by predicting probabilities for subsequent tokens. While training is a computationally intensive process that updates weights on expensive infrastructure, inference utilizes fixed weights and can often be executed on standard consumer hardware.

  • AI model parameters or weights are organized into layers that perform sequential computations.
  • Language models process text by breaking it down into tokens representing words, word fragments, or punctuation.
  • Text generation is an iterative process that repeatedly computes probabilities for and selects the next token.
  • Inference is the process of using an existing trained model, during which weights typically remain unchanged.
  • Model training demands extensive calculations and costly infrastructure to update weights, unlike inference which requires fewer resources.

What Makes it Difficult to Fit the Model on Smaller Hardware

Fitting large language models on smaller hardware is constrained by total memory capacity, memory location, and processing bottlenecks. An 8-billion-parameter model at 16-bit precision requires about 16 GB just for weights, plus extra space for software overhead, temporary calculations, and context. Because persistent storage like SSDs is too slow for execution, weights must reside in fast working memory such as RAM or GPU VRAM. Beyond capacity, memory bandwidth and compute architecture pose limits, with GPUs providing superior parallelism over CPUs and distinct bottlenecks arising across the prefill and decoding stages of inference.

  • Storing an 8-billion-parameter model at 16-bit precision requires approximately 16 GB solely for raw weights, with additional overhead required during execution.
  • System RAM and discrete GPU VRAM remain separate memory spaces with different speeds rather than combining into a single uniform memory pool.
  • Retrieving weights from SSD storage during execution is significantly slower than accessing working memory.
  • GPUs are far more effective than CPUs at performing the parallel numerical operations required by neural networks.
  • Memory bandwidth is often a primary bottleneck during single-user text generation.
  • Inference consists of two distinct stages: parallel prompt processing (prefill) and sequential token generation (decoding).

Why Make the Effort to Run Locally?

Local execution allows developers to utilize existing hardware for experimentation and decrease reliance on rented cloud infrastructure. This approach offers enhanced privacy for sensitive documents and proprietary code, alongside offline capability for desktop assistants and remote deployments. Additionally, running models locally provides greater control over model versions and application behavior. However, it is not universally cost-effective, as hardware, electricity, maintenance, and latency must be weighed against usage frequency compared to hosted services.

  • Running models locally maximizes existing hardware and reduces dependence on rented cloud compute.
  • Local execution safeguards privacy for proprietary source code and private documents.
  • Offline availability is critical for desktop assistants, remote deployments, and intermittent internet scenarios.
  • Developers gain stricter control over model versions and application behavior by hosting locally.
  • Local execution is not inherently cheaper; costs depend on hardware investments, electricity, maintenance, and utilization patterns compared to hosted options.

Quantization: Give Each Weight a Smaller Representation

Quantization reduces the numerical precision used to represent model weights, decreasing storage footprints such as shrinking an 8B model from 16 GB at 16-bit precision to 4 GB at 4-bit precision. While parameter counts stay the same, the lower precision introduces approximation errors that may degrade output quality depending on the task. Lower precision does not guarantee proportional execution speedups, as efficient inference depends on hardware support and runtime conversion software. The technique can be applied after training via post-training quantization or integrated during training to account for approximation errors.

  • Quantization maps numerical weights onto a smaller set of discrete values to reduce memory requirements.
  • Compressing an 8B model from 16-bit to 4-bit precision reduces raw weight storage from approximately 16 GB to 4 GB while keeping parameter count unchanged.
  • Quality degradation from quantization varies based on the specific model, compression method, and target task.
  • Lower precision does not ensure proportional speed improvements because weights may still be unpacked to higher precision during calculation.
  • Quantization can be implemented after training as post-training quantization or accounted for during the training phase itself.

Layer-Wise Offloading: Move Weights as They Are Needed

Layer-wise offloading changes where model weights reside and where their computations execute by moving weights onto the GPU only as needed during a forward pass. While this technique enables the execution of models too large to fit in GPU memory at once, it introduces significant data transfer overhead between RAM, disk, and the GPU. Repeated data movement across generation steps can substantially increase latency, potentially making the model unsuitable for interactive applications. Efficiency can be improved by batching requests or permanently partitioning distinct layers between the CPU and GPU.

  • Quantization alters weight size, whereas offloading changes weight storage location and where calculations execute.
  • Layer-wise offloading transfers layer weights onto the GPU on demand, releasing them before loading subsequent layers.
  • More aggressive offloading strategies store weights on disk and load them into RAM prior to execution.
  • Repeated data transfers across generation steps can severely reduce inference speed, adding notable latency per token.
  • Batching requests to reuse transferred weights can improve throughput when individual latency is not critical.
  • Alternative offloading arrangements statically partition layers across the CPU and GPU or run entirely on the CPU if sufficient RAM is available.

Mixture of Experts: Use Selected Parts for Each Token

Mixture of Experts (MoE) modifies transformer architecture by replacing standard computational blocks with multiple alternative networks known as experts. A routing network dynamically assigns each token to a subset of these experts, separating total parameter count from active parameter count. This allows models to maintain high capacity while reducing per-token computation compared to dense architectures. However, because inactive experts must remain accessible for subsequent tokens, MoE models still demand significant memory storage despite their lower compute footprint.

  • In a sparse MoE model, a routing network dynamically routes each token to a subset of alternative sub-networks called experts.
  • Experts are internal model components whose roles are learned during training, rather than human-interpretable subjects or standalone chatbots.
  • MoE decouples total parameters (which determine storage requirements) from active parameters (which determine per-token computational cost).
  • Inactive experts must remain stored in memory or incur transfer costs when offloaded, meaning memory footprints remain large.
  • MoE reduces per-token computation relative to dense models with equivalent total parameters, but does not inherently reduce the memory requirements needed to run on low-memory devices.

Distillation: Let a Larger Model Teach a Smaller One

Knowledge distillation is a technique where a larger teacher model helps train a smaller student model using its outputs or predictions. The resulting student model functions independently while demanding significantly less memory and computational power for deployment. Distillation produces an entirely distinct model rather than an exact compression, meaning that strong performance on a targeted task does not guarantee broad equivalence across all tasks. The upfront training expenditure allows cost-effective, repeated deployment, though developers can also choose pre-distilled alternatives.

  • Knowledge distillation uses a larger teacher model to train an independent, smaller student model.
  • Student models learn from signals produced by the teacher, such as predictions, enabling them to require less memory and compute.
  • Distillation yields a fundamentally different model rather than an exact, compressed copy of the original teacher.
  • While student models can perform specific tasks well, they may lose breadth, reliability, or the ability to manage unfamiliar problems.
  • The cost of training a distilled model is paid once, facilitating repeated deployments or adoption of pre-distilled models.

Pruning: Remove Work That Contributes Less

Pruning reduces computational overhead by eliminating weights or structural components whose removal results in an acceptable loss of quality. While zeroing individual weights introduces sparsity, it does not automatically reduce storage or execution costs without specialized compressed representations, hardware support, or software engines designed to skip zeroed operations. Alternatively, removing entire units or layers directly downsizes the model structure. Realizing actual efficiency gains requires coordination between the pruned model and the underlying execution runtime, and subsequent retraining may be necessary if too much capacity is removed.

  • Pruning removes weights or larger components that contribute minimally to overall model quality.
  • Zeroing weights creates sparsity, but standard array operations still consume memory and computation unless specifically optimized.
  • Realizing execution savings requires compressed representations, runtime engines capable of skipping inactive calculations, or hardware support.
  • Removing structural components like layers or computation unit groups directly shrinks the model.
  • Excessive pruning degrades model performance, often requiring additional training to recover quality.

The Conversation Has Its Own Memory Requirements

Even after reducing model weights, transformers can encounter out-of-memory errors on long sequences due to dynamic memory requirements. During generation, the key-value (KV) cache stores intermediate representations from prior tokens to avoid redundant calculations. Unlike fixed model weights, the KV cache grows dynamically as sequence length and the number of active conversations increase. Mitigating this memory footprint involves reducing input context, quantizing the cache, or offloading parts of it to CPU memory at the expense of speed or quality.

  • The KV cache stores intermediate token states in transformers to avoid repeated calculations during generation.
  • KV cache memory is separate from model weights and expands with longer context lengths and additional conversations.
  • A static model size does not prevent out-of-memory errors caused by large context workloads.
  • Context size can be managed by retrieving relevant document passages rather than loading full documents.
  • Quantizing the KV cache or offloading it to CPU memory helps manage capacity but can impact generation speed or quality.

Better Software Can Make the Same Hardware More Useful

The implementation of an inference runtime can significantly boost performance on existing hardware without altering model weights. FlashAttention improves execution by partitioning attention calculations into blocks, utilizing fast internal GPU memory to lower memory traffic while computing exact attention. PagedAttention, introduced in vLLM, manages the KV cache in blocks to eliminate memory fragmentation and duplication during dynamic request serving. Additionally, serving software improves throughput by batching multiple requests together to share the overhead of reading model weights, requiring more working memory in exchange.

  • Inference runtime software implementation can substantially impact performance even when model weights remain unchanged.
  • FlashAttention reorganizes attention calculations into blocks to leverage fast GPU memory and reduce memory traffic.
  • FlashAttention computes exact attention without modifying or reducing model weights.
  • PagedAttention was introduced with vLLM to manage the KV cache in blocks, reducing memory waste from inefficient allocation.
  • Serving frameworks batch requests to share the cost of loading model weights, which increases throughput at the cost of higher working memory.

Speculative Decoding: Propose Several Tokens Before Checking Them

Speculative decoding accelerates autoregressive text generation by employing a smaller, faster draft model to propose several candidate tokens. The larger target model evaluates these candidates concurrently, leveraging greater parallelism than standard sequential generation. Accepted tokens are appended to the generation, while rejected tokens are corrected by the target model. Because maintaining a separate draft model increases memory requirements, speculative decoding does not reduce the baseline memory footprint required for the target model.

  • A small draft model proposes multiple candidate tokens ahead of verification.
  • The larger target model evaluates proposed candidate tokens in parallel rather than one by one.
  • Rejected tokens trigger corrections generated directly by the target model.
  • Speed improvements depend on the draft model's agreement rate with the target model and verification costs.
  • Speculative decoding adds memory overhead for the draft model and does not help an oversized target model fit into constrained memory.

Conclusion

Deploying models under hardware constraints, such as a system with 32 GB RAM and 8 GB VRAM, requires balancing memory capacity against response latency and output quality. Applying four-bit quantization reduces an 8B model's theoretical weight footprint from 16 GB to 4 GB, while layer and cache offloading can address remaining memory limits at the expense of transfer speed. A thorough evaluation framework measures answer quality, peak memory consumption, time to first token, and overall generation speed against application-specific demands. Because optimizations like quantization and offloading impact different workload components, their resource savings do not simply compound multiplicatively.

  • Four-bit quantization reduces an 8B model's theoretical weight requirements from 16 GB to 4 GB.
  • Layer or cache offloading can alleviate memory pressure, but the resulting data transfer overhead may degrade generation speed.
  • Model evaluation must track answer quality, peak memory use, time to first token, and generation speed.
  • Tolerance for latency depends heavily on the use case, contrasting interactive code suggestions with background document processing.
  • Optimization savings from quantization, offloading, architectural decisions, and runtime execution do not compound multiplicatively because they target distinct parts of the workload.