Karpathy Said Something You’ll Regret Ignoring
Andrej Karpathy emphasizes that developers remain responsible for software quality even when using vibe coding or agentic engineering. A significant risk in RAG agents is mixed leakage, where a model combines retrieved context with its own parametric knowledge, often leading to ungrounded claims. To address this, the author used Google's Agents CLI and Claude Code to develop a custom evaluation rubric called corpus_abstention. This process identified a specific instruction causing ungrounded answers and improved the agent's performance from 19 of 33 to 30 of 33 by enforcing retrieval. The article examines the necessity of agent-native memory systems that maintain facts across sessions and update dynamically. Research into twelve different systems reveals that reliability is primarily determined at write time, where data structure is established during extraction. Graph-based methods are found to be more effective than flat stores, which often suffer from 'hallucinations of the past' due to stale information. While structured systems are more computationally expensive, they provide a more robust framework that cannot be fixed simply by using larger language models. Large Language Models (LLMs) function by predicting the next word in a sequence using conditional probability based on previous context. During training, the model learns a high-dimensional probability distribution where the trained weights serve as parameters. To prevent repetitive and uncreative outputs, LLMs employ sampling techniques rather than always selecting the most probable token. This process is controlled by a temperature parameter that adjusts the softmax function to balance between deterministic and stochastic generation.
閱讀原文 ↗目錄
Karpathy said something you’ll regret ignoring
Andrej Karpathy emphasizes that developers remain responsible for software quality even when using vibe coding or agentic engineering. A significant risk in RAG agents is mixed leakage, where a model combines retrieved context with its own parametric knowledge, often leading to ungrounded claims. To address this, the author used Google's Agents CLI and Claude Code to develop a custom evaluation rubric called corpus_abstention. This process identified a specific instruction causing ungrounded answers and improved the agent's performance from 19 of 33 to 30 of 33 by enforcing retrieval.
- Developers maintain full responsibility for software vulnerabilities even when using AI agents for vibe coding.
- RAG failures frequently occur when retrieved context is incomplete, leading models to supplement with parametric knowledge.
- Model outputs lack token-level labels to distinguish between retrieved context and internal weights.
- The corpus_abstention rubric provides categorical verdicts to track specific failure modes like mixed leakage.
- Instructions allowing models to skip lookups for simple questions can cause ungrounded responses.
- Implementing mandatory retrieval for every question improved evaluation scores from 19 of 33 to 30 of 33.
Are we ready for an agent-native memory system?
The article examines the necessity of agent-native memory systems that maintain facts across sessions and update dynamically. Research into twelve different systems reveals that reliability is primarily determined at write time, where data structure is established during extraction. Graph-based methods are found to be more effective than flat stores, which often suffer from 'hallucinations of the past' due to stale information. While structured systems are more computationally expensive, they provide a more robust framework that cannot be fixed simply by using larger language models.
- Reliability in agent memory systems is decided at write time, not query time.
- Graph-based methods are more reliable under fact updates and cross-session reasoning.
- Flat stores fail because they lack schemas, leading to generic nodes that hinder downstream filtering.
- Append-only and freeform stores produce 'hallucinations of the past' by returning stale facts.
- Increasing the size of the backbone LLM does not resolve architectural failures in memory pipelines.
- Structured memory systems can be orders of magnitude slower per query than unstructured ones.
- Graphiti is an open-source memory system by Zep that uses Pydantic ontologies for structured extraction.
How do LLMs work?
Large Language Models (LLMs) function by predicting the next word in a sequence using conditional probability based on previous context. During training, the model learns a high-dimensional probability distribution where the trained weights serve as parameters. To prevent repetitive and uncreative outputs, LLMs employ sampling techniques rather than always selecting the most probable token. This process is controlled by a temperature parameter that adjusts the softmax function to balance between deterministic and stochastic generation.
- LLMs predict the next token by calculating the conditional probability P(A|B) for all possible next words given the preceding context.
- The trained weights of an LLM represent the parameters of a high-dimensional probability distribution over sequences.
- Greedy decoding (always picking the highest probability token) leads to repetitive and less useful model outputs.
- Temperature is a hyperparameter that modifies the softmax function to influence the randomness of token sampling.
- Low temperature values result in nearly greedy generation, while high temperature values produce more diverse and random outputs.
- Advanced LLM architectures like Llama 4 utilize components such as Rotary Positional Embeddings (RoPE), sparse routing, and RMSNorm.