← 回到 Reading
ByteByteGo 2026-09-22

How OpenAI Built GPT-Live

Voice systems consistently operate with audio inputs from the user and generate audio responses played back to the user. Although the input and output modalities remain the same, the intermediate processing differs according to the system design. Across their development, voice systems have evolved through three distinct architectural generations. These generations are categorized as cascaded design, turn-based end-to-end, and full-duplex architecture. The cascaded design for voice systems chains three distinct models sequentially: an automatic speech recognition (ASR) model, a large language model (LLM), and a text-to-speech (TTS) model. While this pipeline allows each specialized component to handle what it does best, it introduces significant drawbacks including the loss of vocal cues such as tone and emotion during transcription. Furthermore, operating three systems in series leads to high latency and substantial operational complexity in serving and scaling. These challenges prompted the development of second-generation turn-based speech-to-speech architectures. The second generation of voice AI introduces end-to-end speech-to-speech models that directly process and generate audio, capturing vocal nuances lost in cascaded systems. Despite this improvement, interactions remain turn-based because they still depend on turn-detection and interruption mechanisms to manage speaking flow. These detectors struggle with trade-offs between unnatural delays and premature cutoffs or false triggers. Furthermore, speech-to-speech systems are costly to keep current because integrating updated base LLM checkpoints requires full retraining runs.

閱讀原文 ↗
目錄 15 段
  1. 01Three Generations of Voice Systems
  2. 021. Cascaded Design
  3. 032. Turn-based, Speech-to-Speech
  4. 043. Full-duplex Architecture
  5. 05How the GPT-Live System Works
  6. 06Separating Talking from Thinking
  7. 07The Two Serving Paths
  8. 08How to make the live path fast?
  9. 09The Async Path
  10. 10How to Evaluate a Full-duplex Voice System?
  11. 111. Conversational behavior
  12. 122. The health of the stream
  13. 133. Testing the system on real production traffic
  14. 14What Other Teams Can Learn From GPT-Live
  15. 15What’s Next

Three Generations of Voice Systems

Voice systems consistently operate with audio inputs from the user and generate audio responses played back to the user. Although the input and output modalities remain the same, the intermediate processing differs according to the system design. Across their development, voice systems have evolved through three distinct architectural generations. These generations are categorized as cascaded design, turn-based end-to-end, and full-duplex architecture.

  • The input to a voice system is audio and the output is played back as audio format.
  • The internal processing between audio input and output varies depending on system design.
  • Voice systems have progressed through three distinct architectural generations.
  • The three architectural generations are cascaded design, turn-based end-to-end, and full-duplex architecture.

1. Cascaded Design

The cascaded design for voice systems chains three distinct models sequentially: an automatic speech recognition (ASR) model, a large language model (LLM), and a text-to-speech (TTS) model. While this pipeline allows each specialized component to handle what it does best, it introduces significant drawbacks including the loss of vocal cues such as tone and emotion during transcription. Furthermore, operating three systems in series leads to high latency and substantial operational complexity in serving and scaling. These challenges prompted the development of second-generation turn-based speech-to-speech architectures.

  • A cascaded voice architecture chains an ASR model, an LLM, and a TTS model sequentially.
  • Information loss occurs because the LLM processes text transcripts and misses vocal nuances like tone and emotion.
  • Serial execution causes the latency of all three stages to accumulate, leading to slower user response times.
  • Deploying and managing three distinct models creates significant system complexity in production.
  • The drawbacks of cascaded systems led to the creation of turn-based speech-to-speech models.

2. Turn-based, Speech-to-Speech

The second generation of voice AI introduces end-to-end speech-to-speech models that directly process and generate audio, capturing vocal nuances lost in cascaded systems. Despite this improvement, interactions remain turn-based because they still depend on turn-detection and interruption mechanisms to manage speaking flow. These detectors struggle with trade-offs between unnatural delays and premature cutoffs or false triggers. Furthermore, speech-to-speech systems are costly to keep current because integrating updated base LLM checkpoints requires full retraining runs.

  • End-to-end speech models directly process and output audio, retaining vocal cues lost in cascaded systems.
  • Second-generation speech-to-speech systems remain constrained by turn-based interaction models.
  • Turn detectors face difficult trade-offs between cutting off users prematurely and introducing awkward conversational pauses.
  • Interruption handling requires separate audio-stopping buffers that risk sluggishness or false triggers from background noise.
  • Updating voice models requires expensive full retraining on new LLM checkpoints, causing them to lag behind frontier models.
  • Third-generation full-duplex architectures are designed to overcome the turn-taking problem.

3. Full-duplex Architecture

Full-duplex voice models eliminate traditional turn detectors by continuously listening and generating audio tokens simultaneously, decoding silence as normal tokens. Reference models like Moshi demonstrate this by operating on a fixed clock of roughly 80 milliseconds per frame, learning natural conversational turn-taking directly from training data. While this approach resolves conversation unnaturalness and interruption handling, it introduces high serving costs due to constant inference and demands compact model capacity for millisecond latency. Systems like OpenAI's GPT-Live-1 address these constraints by keeping the primary voice model small and fast while delegating complex reasoning tasks.

  • Full-duplex models eliminate turn detectors by continuously streaming input and output audio tokens, representing silence as an ordinary token.
  • Moshi is an open full-duplex model that converts audio to discrete tokens and operates on a fixed clock of roughly 80 milliseconds per frame.
  • Full-duplex architectures implicitly learn when to speak and when to listen directly from conversational training data.
  • The continuous inference of full-duplex architectures leads to high serving costs and imposes tight constraints on model capacity to maintain millisecond response times.
  • OpenAI launched GPT-Live-1 in July 2026, utilizing a small low-latency voice model that relies on delegation to perform computationally expensive reasoning.

How the GPT-Live System Works

OpenAI developed a voice assistant centered around a full-duplex architecture. The implementation addresses key technical ideas and techniques necessary to operate effectively at scale in production. This approach focuses on making real-time, bidirectional voice interactions viable for large-scale production deployment.

  • OpenAI built a voice assistant structured around a full-duplex architecture.
  • The system focuses on techniques that make real-time voice architectures practical at scale in production.
  • The section examines the engineering methods required to operationalize large-scale full-duplex voice systems.

Separating Talking from Thinking

Standard frontier language models introduce latency when searching and reasoning, which creates awkward silence in voice-based systems. To solve this, OpenAI decouples conversational interaction from deep reasoning tasks. A voice model handles the real-time dialogue, while a separate, more capable model like GPT-5.5 performs necessary reasoning and tool calls in the background. This architecture reduces the traditional trade-off between response latency and answer quality. Furthermore, the modular design simplifies future upgrades by allowing newer reasoning models to be swapped in with minimal engineering effort.

  • OpenAI decouples real-time conversation from complex reasoning to eliminate awkward pauses in voice systems.
  • A dedicated voice model maintains continuous dialogue with the user while backend reasoning tasks run.
  • Complex queries requiring tool use or web searches are delegated to GPT-5.5.
  • Separating conversational duties from reasoning meaningfully reduces the trade-off between speed and answer quality.
  • The modular architecture allows voice assistants to adopt new frontier models with minimal engineering overhead.

The Two Serving Paths

Serving two models concurrently requires decoupling fast and slow operations because audio frames require millisecond latency while delegation can take seconds. To prevent latency spikes, GPT-Live splits traffic into two separate execution pathways. The live path strictly manages rapid audio transfer between the client and voice model. Meanwhile, the async path handles slower, non-audio tasks like tool calls without interfering with live audio streaming.

  • Audio streaming requires frames to be sent every few milliseconds, whereas task delegations can take seconds.
  • A shared process cannot effectively serve both fast audio streaming and slow delegations simultaneously.
  • GPT-Live splits operational traffic into a live path and an async path.
  • The live path exclusively transfers audio between the client and the voice model as quickly as possible.
  • The async path manages non-audio tasks, ensuring slow tool calls do not delay live audio delivery.

How to make the live path fast?

Full-duplex voice systems like GPT-Live require real-time audio transmission on a strict fixed clock to prevent audible artifacts from processing bottlenecks. To minimize initial setup latency, OpenAI created the WARP protocol to compress standard multi-step WebRTC handshakes into a single round trip. Continuous audio inference costs are managed by keeping conversation context pinned in GPU memory on dedicated model instances, alongside batching and speculative decoding. Additionally, OpenAI employs a managed handoff mechanism to seamlessly switch live conversations to replacement instances during context compaction, scaling, or instance failures.

  • Audio in full-duplex systems operates on a fixed clock, meaning any processing bottleneck instantly generates audible artifacts for the user.
  • Standard WebRTC connection setup takes six round trips, introducing over a third of a second of latency on standard 60ms mobile connections.
  • OpenAI built the WARP (WebRTC Abridged Roundtrip Protocol) to parallelize connection steps and initialize sessions in a single round trip.
  • Unlike request-response chatbots, full-duplex systems experience no idle time and sample the model continuously even during user pauses.
  • Inference overhead is reduced by keeping session state loaded directly in GPU memory, allowing the model to process only incoming frames rather than entire conversation histories.
  • A managed handoff mechanism preloads session history onto replacement instances to handle instance updates, failures, and context compaction without interrupting live audio.

The Async Path

The async path is responsible for managing non-audio tasks, including delegations and tool calling, between models in a real-time system. To make the interaction feel seamless, latency must be minimized, particularly the time spent on the prefill phase of prompt processing. OpenAI mitigates this delay by initiating an inference session with the frontier model at the start of a conversation and pre-loading context before a delegation is triggered.

  • The async path manages tasks like tool calling and delegations that take longer than audio processing.
  • Voice models can briefly mask latency through acknowledgments or thinking aloud, but cannot cover prolonged delays.
  • Most delegation latency is caused by prefill, where the model processes the prompt history before generating tokens.
  • In long conversations, prompt processing can add a noticeable fraction of a second to response latency.
  • OpenAI reduces delegation latency by preemptively starting an inference session with the frontier model and sending conversation history in advance.

How to Evaluate a Full-duplex Voice System?

Evaluating full-duplex voice systems requires a different approach than traditional turn-based speech models. Turn-based models can be scored per discrete request-and-response turn, whereas full-duplex architectures operate on one continuous audio stream without distinct turns. Consequently, evaluation in full-duplex architectures focuses on three main components: conversational behavior, stream health, and testing under real production traffic.

  • Turn-based speech models are typically evaluated one turn at a time by scoring discrete responses.
  • Full-duplex architectures operate over a single continuous audio stream without discrete turns, making traditional turn-level evaluation unsuitable.
  • Full-duplex system evaluation focuses on three parts: conversational behavior, the health of the stream, and testing on real production traffic.

1. Conversational behavior

Evaluating conversational behavior in models primarily hinges on timing decisions such as when to speak, remain silent, or interrupt. The two core decisions are endpointing, which determines turn completion, and barge-in detection, which distinguishes genuine interruptions from background noise. These behaviors are evaluated like prediction tasks using annotated natural conversation datasets.

  • Conversational behavior in models primarily depends on appropriate timing.
  • Endpointing involves deciding whether a turn has ended and whether the model should talk or remain silent.
  • Barge-in detection differentiates between actual user interruptions and background noise when the model is speaking.
  • Evaluation treats conversational decisions as predictions scored against annotated natural conversation datasets.

2. The health of the stream

Evaluating stream health in serving systems traditionally relies on percentile targets like p95 latency, which suffices when requests are intermittent. In full-duplex architectures running continuous inference, a p95 event occurs every 20 inferences, resulting in multiple slow events per minute. Because even p99 delays appear regularly, systems must be engineered around p999 targets. In practice, full-duplex designs must emphasize fast recovery since occasional slow frames are inevitable in every session.

  • Traditional serving systems evaluate stream health using percentile targets like p95 latency.
  • In full-duplex continuous inference, p95 latency events happen every 20 inferences, amounting to several times per minute.
  • P99 latency events also occur regularly in continuous streaming architectures.
  • Full-duplex systems must be engineered around p999 latency rather than p95.
  • Systems running continuous streams must be designed for fast recovery because slow frames are inevitable.

3. Testing the system on real production traffic

A silent launch serves as the final layer of evaluation to safely verify system reliability under real production traffic. During this process, users continue interacting with the existing system while a small fraction of voice sessions is routed to the new system. This approach uncovers difficult-to-detect issues, such as unexpected infrastructure bottlenecks. For example, during the silent launch of GPT-Live, the engineering team discovered that a CPU-side service exhausted capacity before the GPUs did.

  • A silent launch safely evaluates system reliability under real production traffic by routing a small percentage of voice sessions to the new system.
  • Users interact with the legacy system as normal while background traffic is evaluated on the new architecture.
  • Silent launches expose edge cases and unexpected bottlenecks that standard synthetic tests fail to detect.
  • In the case of GPT-Live, a silent launch revealed that a CPU-side service reached capacity limits before the GPU infrastructure did.

What Other Teams Can Learn From GPT-Live

Realtime serving architectures, such as OpenAI's GPT-Live, present distinct challenges compared to traditional request-response systems. In voice applications, engineering must prioritize tail latency over average latency, and capacity planning must track concurrent sessions rather than individual requests. Furthermore, moving complex tasks like turn detection directly into the model increases conversational naturalness while reducing overall system complexity. Keeping components outside the model minimal and focused strictly on realtime delivery ensures the broader system remains maintainable.

  • Voice systems require engineering focus on tail latency rather than average latency because users notice the worst-performing frame.
  • Capacity planning for realtime conversational systems must be calculated based on concurrent sessions rather than per-request volume.
  • Absorbing turn detection logic into the model itself improves conversational naturalness and simplifies the serving stack.
  • Non-model components should remain small and dedicated solely to realtime data handling to preserve simplicity and maintainability.

What’s Next

OpenAI views voice-driven computer interaction as the next major paradigm shift in computing, as demonstrated through features in the ChatGPT desktop app like screen perception and asynchronous task execution. Users can continuously query, guide, and redirect tasks via voice in real time. However, challenges persist in smoothly delegating tasks across different models without introducing latency into live conversation. While systems like GPT-Live-1 demonstrate rapid progress likened to the jump from GPT-3 to GPT-4, complex latency and architecture questions remain open.

  • OpenAI envisions voice-driven control of computers as the next generation of human-computer interaction.
  • The ChatGPT desktop app can take screenshots to observe user workflows, run long-running background tasks, and receive mid-task steering via voice.
  • Maintaining low-latency live conversation while delegating complex tasks to auxiliary models remains an active technical challenge.
  • Justin compares the current state of voice AI using GPT-Live-1 to the GPT-4 milestone, having moved past an earlier GPT-3 level of maturity.