← 回到 Reading
Daily Dose of DS 2026-09-11

The Architecture for Serving 100 Fine-Tuned Models on One GPU

This article explores the architectural advantages of serving multiple fine-tuned LLM variants using a shared base model and LoRA adapters instead of merged weights. By utilizing vLLM on Runpod Serverless, the author demonstrates how a shared-base layout significantly reduces GPU memory consumption and improves worker utilization. The experiment compares shared endpoints against separate per-model endpoints, highlighting that shared pools mitigate cold-start delays and operational costs. Ultimately, keeping adapters separate allows a single GPU worker pool to handle a diverse catalog of tasks efficiently.

閱讀原文 ↗
HANDS-ON

The architecture to serve 100 fine-tuned models on a GPU

This article explores the architectural advantages of serving multiple fine-tuned LLM variants using a shared base model and LoRA adapters instead of merged weights. By utilizing vLLM on Runpod Serverless, the author demonstrates how a shared-base layout significantly reduces GPU memory consumption and improves worker utilization. The experiment compares shared endpoints against separate per-model endpoints, highlighting that shared pools mitigate cold-start delays and operational costs. Ultimately, keeping adapters separate allows a single GPU worker pool to handle a diverse catalog of tasks efficiently.

  • Merging 100 fine-tuned 7B model variants requires approximately 1.5 TB of memory, while sharing a base model with LoRA adapters reduces this to 19.3 GB.
  • A rank-8 LoRA adapter contains about 40 MB of weights, whereas a merged model creates a full 15.2 GB copy.
  • Shared endpoints allow for worker reuse across different fine-tuned tasks, preventing the idle capacity and redundant cold starts associated with separate per-variant scaling pools.
  • vLLM can apply requested adapters to a shared base model dynamically, supporting both startup-time and request-time loading for large adapter catalogs.
  • Cold-start latency in serverless environments is primarily driven by image pulling and weight downloading rather than model execution time.
  • In a test using a 1.5B Qwen model, the shared layout maintained consistent performance without the significant request-time penalties of separate endpoints.