← 回到 Reading
Daily Dose of DS 2026-09-29

Build a real-time hotel booking voice agent

Developers have requested the ability to run multiple small models on a single GPU through one inference server, but vLLM has officially declined to implement this feature. The standard workaround requires running separate vLLM instances for each model, which leads to redundant memory and process overhead. An alternative solution called Superlinked Inference Engine (SIE) manages multiple models behind a single server by loading them on demand and using LRU eviction. Performance benchmarks show that SIE can reduce the execution time of a four-model pipeline from 18.58 seconds to 1.47 seconds. This section details the construction of a real-time hotel booking voice agent designed to handle noisy environments and complex data like names and dates. The architecture utilizes a modular pipeline where speech-to-text, reasoning, and speech generation are handled by separate, replaceable components. LiveKit serves as the core infrastructure, integrating Speechmatics Linden for transcription, OpenRouter for LLM access, and Fish Audio for text-to-speech. The implementation emphasizes low latency and the ability to handle mid-sentence corrections through a Streamlit-based monitoring interface. LLM inference memory consumption is divided into model weights, KV cache, activations, and runtime overhead. While weights remain relatively constant, the KV cache scales with context length and concurrent requests, often leading to out-of-memory errors even if the model initially fits. Quantization reduces the weight footprint, which indirectly provides more VRAM for larger batches and longer sequences. Ultimately, inference performance is constrained by memory movement and bandwidth as much as arithmetic throughput.

閱讀原文 ↗
目錄 3 段
  1. 01vLLM closed this feature request as “not planned”:
  2. 02Build a real-time hotel booking voice agent
  3. 03Where does all the VRAM go during LLM inference?
OPEN-SOURCE

vLLM closed this feature request as “not planned”:

Developers have requested the ability to run multiple small models on a single GPU through one inference server, but vLLM has officially declined to implement this feature. The standard workaround requires running separate vLLM instances for each model, which leads to redundant memory and process overhead. An alternative solution called Superlinked Inference Engine (SIE) manages multiple models behind a single server by loading them on demand and using LRU eviction. Performance benchmarks show that SIE can reduce the execution time of a four-model pipeline from 18.58 seconds to 1.47 seconds.

  • vLLM has closed a long-standing feature request for multi-model serving on a single GPU as 'not planned'.
  • The recommended vLLM workaround involves separate processes and memory allocations for every model, requiring an additional routing layer.
  • Superlinked Inference Engine (SIE) allows multiple models to be served through a single inference server on one GPU.
  • SIE manages GPU memory by keeping active models resident and evicting the least-recently-used (LRU) models when space is needed.
  • In testing, SIE improved the performance of a four-model agent pipeline by over 90% compared to direct model loading.
  • SIE is available as an open-source project on GitHub.
HANDS-ON

Build a real-time hotel booking voice agent

This section details the construction of a real-time hotel booking voice agent designed to handle noisy environments and complex data like names and dates. The architecture utilizes a modular pipeline where speech-to-text, reasoning, and speech generation are handled by separate, replaceable components. LiveKit serves as the core infrastructure, integrating Speechmatics Linden for transcription, OpenRouter for LLM access, and Fish Audio for text-to-speech. The implementation emphasizes low latency and the ability to handle mid-sentence corrections through a Streamlit-based monitoring interface.

  • The voice agent pipeline is modular, allowing speech-to-text, reasoning, and speech generation components to be swapped independently.
  • LiveKit Inference simplifies integration by managing authentication, routing, and billing for providers like Speechmatics and Fish Audio.
  • Speechmatics Linden is optimized for voice agents, offering a 250ms latency for final transcripts and high accuracy for alphanumeric strings and accents.
  • The system uses partial transcripts to provide real-time visual feedback in the UI before the final turn is processed.
  • Latency is specifically measured from the completion of a final transcript to the start of the agent's vocal response.
  • The agent's prompt is configured to limit responses to three sentences and prioritize missing reservation fields while preserving user corrections.
LLMs

Where does all the VRAM go during LLM inference?

LLM inference memory consumption is divided into model weights, KV cache, activations, and runtime overhead. While weights remain relatively constant, the KV cache scales with context length and concurrent requests, often leading to out-of-memory errors even if the model initially fits. Quantization reduces the weight footprint, which indirectly provides more VRAM for larger batches and longer sequences. Ultimately, inference performance is constrained by memory movement and bandwidth as much as arithmetic throughput.

  • Model weights are fixed in size based on precision, with FP16 and BF16 requiring two bytes per parameter.
  • The KV cache grows dynamically with context length and the number of concurrent requests.
  • Activations and workspace memory hold intermediate values and vary based on sequence length and batch size.
  • Runtime overhead includes CUDA contexts, memory allocators, and metadata, which are small but essential.
  • Quantization helps by shrinking weights to free up VRAM for larger KV caches or higher concurrency.
  • GPU performance in LLM inference is often limited by memory movement from high-bandwidth memory rather than raw compute throughput.