Do LLMs Have the Memory of a Goldfish?
The term 'memory' in the context of large language models refers to multiple distinct concepts rather than a single unified mechanism. It is critical to differentiate between these meanings to avoid confusion. This section serves as an introduction to examining and clarifying the various types of memory associated with LLMs. During training, large language models encode patterns from vast amounts of data into billions of parameters or weights, enabling general knowledge recall without explicit prompt context. However, standard API interactions do not update these weights, meaning models lack persistent personal memory of individual conversations. Instead, models rely on a temporary context window as working memory to process instructions, conversation history, and retrieved context. Once this context window is cleared, the model loses access to that information unless it is reintroduced by the application. Persistent memory operates external to the model, typically residing in systems such as databases, files, vector stores, or profile services. When required, the application retrieves this stored information and injects it directly into the active context. Consequently, models do not inherently possess persistent memory on their own. Instead, they depend on external auxiliary systems to supply persistent information during execution.
閱讀原文 ↗目錄
- 01What Does “memory” Mean for an LLM?
- 02Trained Memory
- 03Persistent Application Memory
- 04What Happens During a Basic API Call?
- 05Does Every API Call Really Start From Zero?
- 06How AI Chatbots Create the Illusion of Memory
- 07The Context Window is a Working-Memory Budget
- 08Why Long Conversations Become Expensive
- 09What Happens When the Context Window Fills Up?
- 10The Main Techniques to Extend Memory
- 11Sliding-window Memory
- 12Conversation Summarization
- 13Structured Entity Extraction
- 14Vector-store-backed Memory
- 15Long-term User Profiles
- 16How Cross-session Memory Works
- 17Conclusion
What Does “memory” Mean for an LLM?
The term 'memory' in the context of large language models refers to multiple distinct concepts rather than a single unified mechanism. It is critical to differentiate between these meanings to avoid confusion. This section serves as an introduction to examining and clarifying the various types of memory associated with LLMs.
- The word 'memory' has multiple distinct meanings when applied to large language models.
- Different forms of LLM memory should not be conflated with one another.
- A detailed breakdown is required to understand each specific type of LLM memory.
Trained Memory
During training, large language models encode patterns from vast amounts of data into billions of parameters or weights, enabling general knowledge recall without explicit prompt context. However, standard API interactions do not update these weights, meaning models lack persistent personal memory of individual conversations. Instead, models rely on a temporary context window as working memory to process instructions, conversation history, and retrieved context. Once this context window is cleared, the model loses access to that information unless it is reintroduced by the application.
- Patterns learned during LLM training are encoded into billions of numerical parameters or weights.
- Standard API responses do not modify or rewrite model weights, preventing the base model from permanently retaining user-specific facts.
- A model's context window serves as its temporary working memory for generating responses.
- The context window contains system instructions, conversation history, retrieved documents, tool outputs, and user preferences.
- Information in the context window is ephemeral and cannot be recovered by the model once cleared unless re-provided by the application.
Persistent Application Memory
Persistent memory operates external to the model, typically residing in systems such as databases, files, vector stores, or profile services. When required, the application retrieves this stored information and injects it directly into the active context. Consequently, models do not inherently possess persistent memory on their own. Instead, they depend on external auxiliary systems to supply persistent information during execution.
- Persistent memory is maintained outside the model in external systems like databases, vector stores, files, or profile services.
- The application retrieves relevant persistent information and inserts it into the model context when needed.
- Models do not natively possess persistent memory.
- All persistent information utilized by the model is delivered to it by external systems.
What Happens During a Basic API Call?
AI model API calls are inherently stateless, meaning the model does not store an internal record or memory of prior interactions once a request finishes. If an application requires context from earlier turns, it must re-send the preceding conversation history within the current request payload. As a result, most providers implement a stateless Messages API that depends on the client supplying the full dialogue history for multi-turn conversations.
- AI models do not retain memory or state across independent API calls.
- A request lacking previous context cannot reliably access details supplied in earlier requests.
- Applications must retransmit previous exchanges in multi-turn dialogues to preserve context.
- Most model providers utilize a stateless Messages API design that expects the conversation history in each request.
Does Every API Call Really Start From Zero?
A new model API call does not start completely from zero because it retains its trained weights, general language capabilities, platform safety instructions, and provided context. However, models inherently lack an automatically updated personal memory of specific users or past conversations. Outputs are strictly generated from the model's static weights combined with whatever context is supplied for that specific invocation. Even when APIs offer server-managed conversation state via session identifiers, the server is simply reconstructing the context for the model rather than imparting personal memory to it.
- A model invocation retains trained weights, language abilities, training knowledge, and platform safety instructions.
- Models do not possess an automatically updated memory of specific users or ongoing conversations.
- Every output is generated purely from the model's weights and the context supplied for that specific invocation.
- Prior conversation history is completely unavailable unless the surrounding system explicitly forwards it.
- Server-managed conversation state uses conversation identifiers to reconstruct context behind the scenes without changing the model's stateless nature.
How AI Chatbots Create the Illusion of Memory
AI chatbots create the appearance of conversational recall by appending the full dialogue history, including system instructions and past user-assistant turns, into each subsequent prompt. To the underlying model, this interaction is simply an evaluation of a single large input rather than true temporal retention. Because application-level continuity exists, this technique is more appropriately classified as reconstructed or context-based memory rather than fake memory.
- Chatbot applications construct a single large input containing all previous conversation turns and system instructions for each new turn.
- Underlying language models do not store persistent memories across interactions, but process prior conversational context present in the active prompt.
- Continuity in chat systems exists at the application level rather than within the model itself.
- The phenomenon of passing conversation history as prompt context is best described as reconstructed memory or context-based memory.
The Context Window is a Working-Memory Budget
The context window of a large language model acts as a finite working-memory budget measured in tokens. Beyond the visible conversation, this budget must accommodate system instructions, tool definitions, retrieved documents, and the output response. When capacity is reached, applications must remove, compress, or replace content to stay within limits. Crucially, larger context windows do not ensure perfect recall, as models face a degradation phenomenon known as context rot when processing excessive information.
- Context windows are measured in tokens, which represent words or parts of words.
- A context window must allocate space for hidden system instructions, tool definitions, retrieved documents, and generated output in addition to visible conversation history.
- Once context limits are met, an application must compress, remove, or replace existing information.
- Larger context windows do not guarantee accurate recall due to a degradation problem known as context rot.
- Managing the context window requires active attention management and information filtering, not just raw token capacity.
Why Long Conversations Become Expensive
LLM APIs charge for processed tokens, causing the cumulative token cost of long conversations to grow rapidly as earlier history is reprocessed with every new turn. In addition to escalating financial costs, longer contexts increase generation latency. Prompt caching can mitigate latency and repetitive processing expenses by reusing prefixes such as system prompts and history. However, because cached tokens still consume context window space, prompt caching serves as an efficiency optimization rather than a true memory architecture.
- Conversational rounds reprocess prior history on every request, leading to cumulative input token growth that exceeds the visible conversation size.
- Longer contexts increase latency because the model must process more information before beginning output generation.
- Additional elements like system prompts, tool definitions, and retrieved documents compound per-request token consumption.
- Prompt caching reduces cost and latency by reusing precomputed prefixes like system prompts and conversation history.
- Prompt caching is an optimization technique and does not grant unlimited memory or expand the context window.
What Happens When the Context Window Fills Up?
Large language models and their hosting applications have no single universal behavior when reaching context window limits. Instead, systems use various strategies such as request rejection, message truncation, rolling windows, and compaction. Compaction preserves the core state of a conversation in a reduced representation to enable long-running sessions. However, this summarization technique is inherently lossy, meaning seemingly unimportant details might be prematurely discarded.
- Context window exhaustion has no universal standard behavior across LLM providers and applications.
- Common strategies for handling full context windows include rejecting requests, evicting oldest messages, using rolling windows, and substituting past context with compact summaries.
- Compaction preserves essential conversation state in a smaller format to permit extended interactions under reduced context usage.
- Summarization in context management is lossy, creating a tradeoff where information deemed low-value initially may be needed later.
The Main Techniques to Extend Memory
Real-world applications typically require combining multiple approaches to extend memory rather than relying on an isolated solution. Integrating different techniques allows systems to handle memory constraints more effectively in practice. This section introduces an exploration into several of these specific memory-extension methods.
- Practical applications generally combine multiple techniques to extend memory.
- Relying on a single memory extension technique is usually insufficient for real-world systems.
- A variety of distinct techniques exist to address memory extension in detail.
Sliding-window Memory
Sliding-window memory retains only the most recent portion of a conversation by discarding the oldest messages when new ones arrive. A typical setup preserves system instructions, a fixed number of recent turns, and the current user input. While simple, fast, and predictable, this technique is best suited for dialogues where recent context matters far more than historical context. Its main disadvantage is permanent information loss once earlier requirements or facts exit the active context window.
- A sliding window retains the newest messages and removes the oldest as conversations progress.
- Common setups preserve system instructions, the current user message, and a fixed number of recent turns (e.g., 20 turns).
- The technique provides simplicity, fast execution, and predictable memory usage.
- Effectiveness relies on recent conversational context being more critical than past messages.
- Facts that leave the context window are forgotten, preventing the model from recalling earlier constraints.
Conversation Summarization
Conversation summarization periodically compresses older conversation turns into concise descriptions to maintain the context efficiently. A standard context structure consists of system instructions, the conversation summary, recent unsummarized messages, and the current user query. However, summarization acts as an interpretation that strips nuance, simplifies uncertainty, and risks converting assumptions into false facts. To avoid progressive information distortion from repeated summarization, essential facts should be stored separately rather than entrusted to a narrative summary alone.
- Summarization periodically condenses older messages into a compact description while discarding original turns.
- A typical conversation context layout combines system instructions, the summary, recent unsummarized turns, and the current user prompt.
- Repeatedly summarizing prior summaries causes cumulative distortion, similar to photocopying a photocopy.
- Summaries tend to lose nuance, flatten uncertainty, and can misrepresent assumptions as confirmed facts.
- Critical facts are more reliably preserved by storing them separately rather than relying on iterative narrative summaries.
Structured Entity Extraction
Structured entity extraction records specific conversational facts into predefined fields rather than storing interactions as raw prose. This approach provides higher reliability when an application needs to access exact project state and confirmed parameters. While effective for discrete items like user preferences, technical decisions, and workflow statuses, it struggles to capture nuanced narrative details that resist structured schemas.
- Structured entity extraction parses specific facts into structured fields instead of relying on prose conversation logs.
- Querying structured fields is more reliable for exact project state retrieval than searching long text summaries.
- Structured memory is well-suited for discrete data including user preferences, technical decisions, workflow statuses, and deadlines.
- Subtle, narrative information that cannot be mapped to predefined schemas is poorly served by structured entity extraction.
Vector-store-backed Memory
Vector-store-backed memory enables conversational systems to perform semantic retrieval rather than appending entire conversation histories into context. The system segments past dialogue into chunks, transforms them into embeddings, and compares them against incoming queries to identify relevant passages. Only the highest-ranking passages are inserted into the prompt, effectively managing the model's context window. However, semantic search can retrieve irrelevant, outdated, or poorly matched memories, making metadata filtering crucial for accuracy.
- Vector stores enable semantic retrieval by segmenting past conversations and converting them into embeddings.
- Embeddings numerically capture meaning, allowing matches between texts that express similar concepts without identical wording.
- The workflow embeds new queries, retrieves semantically similar memory candidates, filters and ranks them, and inserts the top results into the prompt.
- Semantic retrieval selects relevant memories for the context window rather than increasing the window size.
- Failure modes include retrieving irrelevant matches, missing memories phrased unusually, and returning outdated decisions.
- Using metadata like user ID, project ID, date, and memory type is essential to prevent retrieval errors.
Long-term User Profiles
A long-term user profile stores durable preferences and facts that remain useful across multiple conversations. Systems must distinguish between stable, repeatedly confirmed preferences and transient session-level details. Profiles require ongoing updates because user preferences evolve over time. Effective implementations enrich memory entries with metadata such as timestamps, sources, and confidence scores.
- Long-term user profiles preserve durable facts across conversations rather than transient statements.
- Session information reflects temporary activities, whereas durable preferences require repeated confirmation.
- User profiles must be updated over time to accommodate evolving preferences.
- Robust memory systems tag memory entries with metadata including timestamps, sources, and confidence scores.
How Cross-session Memory Works
Cross-session memory enables information to persist across separate chat sessions and influence subsequent interactions. A background process identifies durable details from a conversation and saves them into a user-, project-, or organization-level store. When a new session begins, relevant memories are retrieved and explicitly inserted into the model's active context window. This context insertion step is critical, as the model requires the remembered information to be present in its prompt to inform responses.
- Cross-session memory allows information to survive past the end of an individual chat session.
- Durable information is identified and saved to a user-level, project-level, or organization-level memory store.
- For cross-session memory to function, retrieved memories must be explicitly injected into the new conversation's context.
- The underlying model generates responses relying on the supplied memories placed within its current context.
Conclusion
Large language models do not intrinsically carry personal experiences or episodic memory from one interaction to the next. Instead, apparent LLM memory is produced by an architectural system combining model weights, context windows, conversation storage, long-term memory, and a memory manager. This setup relies on processes including context reconstruction, summarization, retrieval, and persistent storage. Ultimately, the quality and effectiveness of an LLM application's memory depend heavily on its surrounding software architecture rather than just the underlying model.
- Model weights contain general knowledge acquired during training but do not retain interaction history across calls.
- The context window holds only the information active for generating the current response.
- Conversation storage preserves previous messages, whereas long-term memory stores selected facts, preferences, and past events.
- The memory manager governs what information is retrieved and placed into active context.
- LLMs do not automatically carry personal experiences from one API call to the next.
- LLM memory is an engineered system of context reconstruction, summarization, retrieval, and persistent storage.