The Architecture for Serving 100 Fine-Tuned Models on One GPU
This article explores the architectural advantages of serving multiple fine-tuned LLM variants using a shared base model and LoRA adapters instead of merged weights. By utilizing vLLM on Runpod Serverless, the author demonstrates how a shared-base layout significantly reduces GPU memory consumption and improves worker utilization. The experiment compares shared endpoints against separate per-model endpoints, highlighting that shared pools mitigate cold-start delays and operational costs. Ultimately, keeping adapters separate allows a single GPU worker pool to handle a diverse catalog of tasks efficiently.
閱讀原文 ↗The architecture to serve 100 fine-tuned models on a GPU
This article explores the architectural advantages of serving multiple fine-tuned LLM variants using a shared base model and LoRA adapters instead of merged weights. By utilizing vLLM on Runpod Serverless, the author demonstrates how a shared-base layout significantly reduces GPU memory consumption and improves worker utilization. The experiment compares shared endpoints against separate per-model endpoints, highlighting that shared pools mitigate cold-start delays and operational costs. Ultimately, keeping adapters separate allows a single GPU worker pool to handle a diverse catalog of tasks efficiently.
- Merging 100 fine-tuned 7B model variants requires approximately 1.5 TB of memory, while sharing a base model with LoRA adapters reduces this to 19.3 GB.
- A rank-8 LoRA adapter contains about 40 MB of weights, whereas a merged model creates a full 15.2 GB copy.
- Shared endpoints allow for worker reuse across different fine-tuned tasks, preventing the idle capacity and redundant cold starts associated with separate per-variant scaling pools.
- vLLM can apply requested adapters to a shared base model dynamically, supporting both startup-time and request-time loading for large adapter catalogs.
- Cold-start latency in serverless environments is primarily driven by image pulling and weight downloading rather than model execution time.
- In a test using a 1.5B Qwen model, the shared layout maintained consistent performance without the significant request-time penalties of separate endpoints.