← 回到 Reading
ByteByteGo 2026-08-31

What Happens Inside an AI Chatbot Between Enter and the First Word?

User queries are not sent directly to large language models in isolation; instead, an assembled document containing system prompts, tool definitions, memory, retrieved documents, and conversation history is provided. Managing what goes into this input and how it is structured is an essential discipline known as context engineering. Because models possess finite attention budgets, longer inputs can gradually degrade reasoning precision even on basic tasks. Consequently, products using the identical underlying model can generate differing responses depending on how they curate and time the assembly of context. Large language models are stateless by design, requiring the entire conversation history to be resent with every new turn. Because conversation history accumulates with each interaction, input token volume compounds and typically dominates total API costs despite lower per-token pricing than outputs. When exchanges exceed the context window, naive retransmission breaks down, necessitating strategies such as dropping older turns, summarizing chat history, or using external storage and retrieval. Consequently, long conversations suffer from escalating latency, higher expenses, and potential loss of fine-grained conversational details. AI systems can perform input safety checks using a distinct, decoupled safety model before answer generation begins. This architectural separation allows the safety mechanism to be independently retrained, monitored, and used for routing or escalations without modifying the primary assistant. While early single-classifier designs incurred an expensive 24 percent compute overhead and raised false refusals by 0.38 percentage points, modern deployments resolve this via a cascade architecture. In this cascaded setup, a low-cost validation inspects all traffic, routing only flagged inputs to the more resource-intensive classifier to bring compute overhead down to approximately one percent.

閱讀原文 ↗
目錄 10 段
  1. 01How the Input to the Model is Assembled?
  2. 02Why is Every Input Message to the Model Independent?
  3. 03Performing Safety Checks on the Input
  4. 04How the Model Understands the Words?
  5. 05How the Model is Shared Across Multiple Conversations
  6. 06Prefill And Decode
  7. 07Caching the Existing Calculations
  8. 08Streaming And Guardrails
  9. 09How Tools Are Run?
  10. 10Conclusion

How the Input to the Model is Assembled?

User queries are not sent directly to large language models in isolation; instead, an assembled document containing system prompts, tool definitions, memory, retrieved documents, and conversation history is provided. Managing what goes into this input and how it is structured is an essential discipline known as context engineering. Because models possess finite attention budgets, longer inputs can gradually degrade reasoning precision even on basic tasks. Consequently, products using the identical underlying model can generate differing responses depending on how they curate and time the assembly of context.

  • The actual input sent to a model is a curated document consisting of system prompts, tool definitions, memory, retrieved data, conversation history, and the user prompt.
  • Context engineering is the discipline of selecting, structuring, and pruning material within the model's input document.
  • LLMs have a finite attention budget, meaning that precision and accuracy degrade gradually as input length grows.
  • Two products powered by the identical model can produce different responses to identical queries due to differences in context engineering.
  • Context material can be retrieved completely up front for faster speed or retrieved dynamically via lightweight references to conserve tokens.

Why is Every Input Message to the Model Independent?

Large language models are stateless by design, requiring the entire conversation history to be resent with every new turn. Because conversation history accumulates with each interaction, input token volume compounds and typically dominates total API costs despite lower per-token pricing than outputs. When exchanges exceed the context window, naive retransmission breaks down, necessitating strategies such as dropping older turns, summarizing chat history, or using external storage and retrieval. Consequently, long conversations suffer from escalating latency, higher expenses, and potential loss of fine-grained conversational details.

  • Models retain no memory between requests, requiring the full conversation history to be reconstructed and resent on every turn.
  • Input token counts compound with each exchange, causing input costs to dominate overall spend despite lower unit pricing than output tokens.
  • The naive strategy of resending the entire conversation fails once the interaction exceeds the model's context window.
  • Context mitigation strategies include trimming older turns, summarizing prior dialogue, or storing data externally for on-demand retrieval.
  • Extended interactions result in increased processing latency, higher costs, and potential information loss due to trimming or condensation.

Performing Safety Checks on the Input

AI systems can perform input safety checks using a distinct, decoupled safety model before answer generation begins. This architectural separation allows the safety mechanism to be independently retrained, monitored, and used for routing or escalations without modifying the primary assistant. While early single-classifier designs incurred an expensive 24 percent compute overhead and raised false refusals by 0.38 percentage points, modern deployments resolve this via a cascade architecture. In this cascaded setup, a low-cost validation inspects all traffic, routing only flagged inputs to the more resource-intensive classifier to bring compute overhead down to approximately one percent.

  • Input safety checks utilize a decoupled, smaller model to evaluate requests before the primary assistant generates a response.
  • Decoupling the safety layer permits independent retraining, fine-tuning, monitoring, and actions such as logging, routing, or human escalation.
  • Early production safety classifiers increased compute requirements by approximately 24 percent and increased harmless query refusals by 0.38 percentage points.
  • A cascaded safety architecture deploys a lightweight validator across all traffic and invokes the resource-intensive classifier only on flagged inputs.
  • Cascaded validation cuts compute overhead down to roughly one percent and harmless query refusal rates to 0.05 percent.

How the Model Understands the Words?

Documents are converted into tokens, sub-word chunks of text that models process as their primary units. Modern systems typically construct tokens from raw bytes to ensure representation across all writing systems, averaging roughly three-quarters of an English word per token. Because token counts for the same content can vary across translations by up to fifteen times, speakers of certain languages face higher costs, slower processing, and reduced effective context window capacity. Furthermore, chunking text into tokens makes fine-grained, character-level tasks like counting specific letters awkward for models.

  • Tokens are sub-word text chunks positioned between individual characters and whole words, averaging approximately 0.75 English words each.
  • Modern systems build tokens from raw bytes rather than characters to reliably represent any writing system.
  • Equivalent text translated into different languages can exhibit token count variations of up to 15 times.
  • Higher token density per text reduces usable context window space, increases latency, and inflates financial cost.
  • Operating on multi-character token chunks complicates character-level reasoning tasks, such as letter counting.

How the Model is Shared Across Multiple Conversations

Serving large models economically requires batching multiple requests together to minimize memory bandwidth bottlenecks caused by repeatedly loading model parameters. While naive batching underutilizes hardware due to differing output lengths among conversational requests, scheduling at individual generation steps allows new requests to fill finished slots immediately. This step-level scheduling approach can achieve up to 23 times higher throughput and lower median response times over naive batching. However, concurrent execution causes numerical variations, meaning identical prompts can produce different outputs even when randomness is disabled.

  • Batching multiple requests amortizes the memory bandwidth cost of loading model parameters on modern accelerators.
  • Naive batching leaves hardware partially idle because requests have variable lengths, forcing compute to wait for the longest completion.
  • Scheduling requests at the level of individual generation steps can improve throughput by up to 23 times over naive batching while also lowering median response times.
  • Sharing compute hardware across varying batch sizes causes numerical variations, leading identical prompts to produce dozens of distinct outputs even with randomness turned off.

Prefill And Decode

Generating a response occurs in two distinct phases: an initial prefill phase that processes all input tokens in parallel and a subsequent decode phase that generates output tokens sequentially. The prefill phase is computationally bound and scales with input length, leading to longer wait times before output begins in extended conversations. In contrast, the decode phase is memory-speed bound and outputs tokens at a relatively steady rate regardless of input size. System performance is measured through time to first token and time per output token, with techniques like input chunking used to prevent long requests from stalling concurrent batch generation.

  • Inference is split into two phases: a parallel, compute-bound prefill phase and a sequential, memory-bound decode phase.
  • The prefill phase duration scales with the length of the input context, causing the pause before generation starts to increase in longer conversations.
  • The decode phase produces tokens at a roughly constant rate regardless of the prior conversation length.
  • Total response time approximately equals the time to first token plus the time per output token multiplied by reply length.
  • Splitting long inputs into chunks and interleaving them with ongoing token generation prevents batched requests from stalling.

Caching the Existing Calculations

During the generation phase of language model inference, intermediate calculations from previous tokens are stored in memory and reused to avoid repeating redundant computation. This conversation state consumes significant memory, often making it the primary bottleneck for serving capacity rather than the model weights themselves. Borrowing virtual memory techniques from operating systems, modern serving architectures allocate storage in small on-demand blocks, cutting memory waste from up to 80% to under 4% and increasing throughput by 2-4 times. Additionally, caching reusable prompt prefixes across turns provides substantial cost discounts and motivates structuring prompts with static content placed before dynamic content.

  • Intermediate token calculations are cached and reused across inference steps, but can consume gigabytes of memory per request for a 70-billion-parameter model with an 8000-token context.
  • Concurrent user serving capacity is typically limited by the memory required for conversation state rather than model weights.
  • Early serving systems reserved contiguous memory blocks for the maximum possible response length, leaving 60% to 80% of memory unused.
  • Allocating memory in small fixed-size blocks on demand reduces memory waste to under 4% and improves serving throughput by 2 to 4 times.
  • Cached portions of prompts are often priced around one-tenth of normal input rates by API providers.
  • Prompt caching relies on prefix reuse, necessitating that stable content be placed at the top and changing content at the bottom of prompts.

Streaming And Guardrails

During the AI response phase, streaming delivers text word-by-word to match or exceed human reading speeds of roughly six tokens per second, substantially improving perceived latency. However, conventional output safety checks require evaluating the completed response before judging it, making it impossible to retract already displayed words. Delaying output until safety verification finishes eliminates the benefits of streaming. Emerging alternatives to address this include token-by-token stream evaluation and inspecting internal model states during generation.

  • Streaming output word-by-word improves perceived generation speed by avoiding a blank waiting period.
  • Human reading speed averages around six tokens per second, which production systems comfortably exceed.
  • Traditional output safety checks inspect entire finished responses, conflicting with the real-time nature of streaming.
  • Buffering responses until safety checks complete entirely negates the perceived speed advantage of streaming.
  • Modern workarounds include guards that inspect text token-by-token or inspect internal model states during generation.

How Tools Are Run?

Integrating tools changes the standard linear generation process into an iterative execution loop. Language models do not execute external actions like web searches directly; rather, they emit text requesting an action, which the host application fulfills and appends back into context. Each completed tool action restarts the full pipeline—including reassembly, tokenization, queuing, and generation. Consequently, multi-step queries incur noticeable latency and sharply compounding token costs as accumulating context is repeatedly processed across successive calls.

  • Models do not directly execute tools like web searches or database queries; they output formatted text requests for the surrounding application to run.
  • Tool execution converts the one-way generation path into a multi-step loop where results are fed back into context as new input.
  • Each tool call forces the entire pipeline to rerun from scratch, including assembly, checks, tokenization, queuing, and two generation phases.
  • Multiple round trips across tools create noticeable latency compared to simple generation.
  • Context costs compound severely because early messages and instructions are resent and billed on every subsequent call within a session.

Conclusion

In an LLM interaction, the end-to-end response delay is driven by multiple components including network latency, document assembly, safety checks, queueing, and input reading. Because language models retain no inherent memory, conversational history must be reconstructed and re-read from scratch with every turn. In extended conversations, processing this accumulated input context consumes the largest fraction of initial latency. Consequently, longer chats suffer from increased delay, higher operational costs, and potential detail loss.

  • Input reading accounts for the largest fraction of initial delay in extended conversations, causing the initial pause to lengthen over time.
  • Models have no stateful memory, requiring the full conversation context to be rebuilt from scratch on every turn.
  • Input safety checks utilize cascaded designs that account for approximately one percent of total compute.
  • The overhead incurred from tokenization is negligible in overall response latency.
  • Document assembly and context gathering constitute a larger fraction of delay than raw network overhead and authentication.