← 回到 Reading
ByteByteGo 2026-08-25

How to Steal an AI Model’s Private Thoughts

Modern frontier models produce an extended internal sequence of text prior to outputting a visible response, allowing them to explore hypotheses and correct errors. This internal sequence is referred to as a reasoning trace or chain of thought. Unlike the final polished output, the reasoning trace holds raw tool outputs, intermediate exploration, and any sensitive contextual data processed during the session. For instance, when a coding agent is instructed to purge hardcoded credentials, those credentials inevitably pass through the reasoning trace. Model providers conceal reasoning traces primarily for commercial and safety-related reasons. Commercially, competitors can harvest traces from a capable model to extract its methodology and train cheaper imitation models. From a safety standpoint, models often generate intermediate reasoning about harmful subjects prior to issuing a refusal. Because safety filtering is applied after trace generation to produce the final response, releasing raw traces bypasses these guardrails and exposes harmful material to users. In multi-turn conversations, earlier reasoning traces must remain accessible across requests despite model APIs being fundamentally stateless. Providers face two architectures: storing traces on the server referenced by session identifiers, or encrypting traces and delegating storage back to the client. Leading providers including OpenAI, Anthropic, and Google chose the client-side encrypted approach. This architecture satisfies confidentiality and integrity while eliminating the infrastructure cost of server-side trace storage.

閱讀原文 ↗
目錄 10 段
  1. 01Reasoning Traces
  2. 02Concealment Rationale
  3. 03State Management
  4. 04Envelope Structure
  5. 05Trace Compatibility
  6. 06Extraction Method
  7. 07Attack Vectors
  8. 08Field Observations
  9. 09Proposed Mitigations
  10. 10Conclusion

Reasoning Traces

Modern frontier models produce an extended internal sequence of text prior to outputting a visible response, allowing them to explore hypotheses and correct errors. This internal sequence is referred to as a reasoning trace or chain of thought. Unlike the final polished output, the reasoning trace holds raw tool outputs, intermediate exploration, and any sensitive contextual data processed during the session. For instance, when a coding agent is instructed to purge hardcoded credentials, those credentials inevitably pass through the reasoning trace.

  • Frontier models generate extended internal text sequences before producing a final answer.
  • A reasoning trace is also known as a chain of thought.
  • Reasoning traces capture intermediate hypotheses, abandoned paths, and raw tool outputs.
  • Reasoning traces are denser and more revealing than final outputs, often holding contextual secrets or user data.
  • Tasks such as removing credentials cause sensitive data to be exposed within the reasoning trace.

Concealment Rationale

Model providers conceal reasoning traces primarily for commercial and safety-related reasons. Commercially, competitors can harvest traces from a capable model to extract its methodology and train cheaper imitation models. From a safety standpoint, models often generate intermediate reasoning about harmful subjects prior to issuing a refusal. Because safety filtering is applied after trace generation to produce the final response, releasing raw traces bypasses these guardrails and exposes harmful material to users.

  • Model providers hide reasoning traces due to commercial competition and safety concerns.
  • Exposing reasoning traces reveals the methodology behind computations, allowing competitors to train cheaper imitation models.
  • Models frequently generate reasoning about harmful topics during the process of formulating a refusal.
  • Safety filtering operates downstream on completed traces to generate safe visible answers.
  • Publishing reasoning traces circumvents post-generation safety filters, potentially exposing users to harmful content.

State Management

In multi-turn conversations, earlier reasoning traces must remain accessible across requests despite model APIs being fundamentally stateless. Providers face two architectures: storing traces on the server referenced by session identifiers, or encrypting traces and delegating storage back to the client. Leading providers including OpenAI, Anthropic, and Google chose the client-side encrypted approach. This architecture satisfies confidentiality and integrity while eliminating the infrastructure cost of server-side trace storage.

  • Model APIs are inherently stateless, meaning conversation continuity must be explicitly reconstructed with each request.
  • Server-side state storage tracks conversation traces in a database using identifiers, which is straightforward but expensive at global scale.
  • Client-side state storage encrypts the trace and returns it to the client, requiring the provider to store nothing between turns.
  • OpenAI, Anthropic, and Google all use the client-side encrypted trace approach.
  • The client-side encrypted approach achieves confidentiality against competitors, integrity against trace tampering, and cost savings via statelessness.

Envelope Structure

The payload block returned to clients is formatted as a base64-encoded AEAD (Authenticated Encryption with Associated Data) envelope that provides both encryption and tamper protection. The envelope consists of a header with metadata like model name and version, along with a nonce, authentication tag, and ciphertext. Different providers name this container differently, such as Anthropic using signature, OpenAI using encrypted_content, and Google using thinkingSignature. Notably, authentication covers the model name and version but omits account and conversation identifiers, with observed behavior suggesting providers use a single global key across an ecosystem.

  • Blocks returned to clients are base64-encoded AEAD envelopes that conceal content while ensuring tamper-protection.
  • The envelope contains a header (e.g., model name, block type, version, key ID), a nonce, an authentication tag, and ciphertext.
  • The payload field is named 'signature' by Anthropic, 'encrypted_content' by OpenAI, and 'thinkingSignature' by Google.
  • Authentication covers the model name and version, but excludes user account and conversation identifiers.
  • Observable behavior suggests providers use a single global key across their entire ecosystem.

Trace Compatibility

Researchers analyzed block compatibility when authenticated fields lack origin information, identifying three progressively permissive forms: cross-session, cross-user, and cross-model compatibility. Cross-model compatibility allows for dynamic tasks like mid-conversation model switching and automatic rerouting. An evaluation conducted in July 2026 across various model combinations revealed differing levels of support across major AI models. While Gemini accepted all combinations across generations and Claude accepted nearly all combinations except Fable 5, GPT's compatibility was generation-dependent, with GPT-5.6 accepting blocks from older generations.

  • Block compatibility is categorized into three levels: cross-session, cross-user, and cross-model compatibility.
  • Cross-session compatibility allows replaying blocks out of order, modifying conversation history, and trimming sessions to context limits.
  • Cross-model compatibility permits blocks from one model to be accepted by another, enabling mid-conversation model switching and rerouting.
  • Gemini accepted every tested block combination across every generation in tests from July 2026.
  • Claude accepted almost all combinations, but Fable 5 blocks were only accepted by Fable 5 itself.
  • The GPT-5.6 series accepted blocks from all earlier GPT generations, while older GPT models only accepted their own blocks.

Extraction Method

Flagship AI models undergo anti-distillation training to conceal their reasoning, but smaller models within the same family receive significantly less protection due to cost and speed optimizations. Researchers leveraged this cross-model compatibility gap by extracting encrypted reasoning blocks from flagship models and prompting smaller models to transcribe them into plaintext. The fidelity of the recovered reasoning was verified by matching re-encoded token counts against exact reasoning token counts in API billing records. Extraction difficulty differed between providers, requiring a single prompt for Claude while GPT required chunking and dozens of candidate extractions.

  • Flagship models receive anti-distillation training that smaller, cost-optimized models in the same family largely lack.
  • Reasoning traces from strong models can be recovered by feeding their reasoning blocks into weaker sibling models acting as fuzzy decoders.
  • The flagship model's refusal training is bypassed because it is only queried with an ordinary question that generates a benign output.
  • Billing records reporting exact reasoning token counts were used across 120 programming tasks to verify reconstruction faithfulness.
  • Claude models required only a single prompt for extraction, whereas GPT models required chunked outputs and up to 50 attempts per block to evade verbatim reproduction filters.

Attack Vectors

The research outlines four distinct attack vectors arising from hidden or encrypted model reasoning traces, categorized by whether the blocks are created by attackers or third parties. Attackers can use reasoning traces to enhance model distillation and bypass visible safety guards through jailbreaking hidden reasoning. Furthermore, third-party agent logs shared on platforms like GitHub and Hugging Face routinely leak un-sanitized secrets including API keys and passwords. Finally, attackers can plant malicious instructions inside encrypted blocks of shared agent trajectories to execute prompt injection attacks upon session resumption.

  • Distillation leveraging reasoning traces raised MATH500 accuracy from 68.4 percent to 76.0 percent compared to answer-only training.
  • Decoding 10,000 reasoning traces costs roughly $720 at Claude Haiku 4.5 pricing.
  • Jailbreaking models via traces is viable because models are rarely trained to constrain harmful reasoning, allowing sensitive details to stay hidden inside traces while visible outputs appear benign.
  • An analysis of 6,708 public trajectories across GitHub and Hugging Face revealed 1,028 leaking blocks containing 62 API keys, 33 passwords, and 24 access tokens.
  • Encrypted blocks in resumed agent runs facilitate prompt injection, demonstrated by tricking GPT-5.6 Sol into exfiltrating PowerPoint files.

Field Observations

Recovering reasoning traces at scale reveals that returned summaries can diverge substantially from actual internal model reasoning. In an example from an AIME 2025 problem, an API summary claimed methodical geometric derivation, whereas the underlying trace showed memory recall and trial-and-error backward justification. Additionally, recovered GPT traces typically consist of compressed, ungrammatical telegraphic notes rather than fluent English. This lack of legibility creates significant hurdles for trace monitoring and AI oversight.

  • API-generated summaries sometimes diverge from the actual reasoning path observed in recovered traces.
  • On an AIME 2025 problem, the summary reported methodical geometry steps using the law of cosines, while the trace showed post-hoc rationalization from memory.
  • Smaller models tasked with summarizing larger models inherently compress and filter the original thought process.
  • Recovered GPT reasoning traces often consist of telegraphic, ungrammatical notes with dropped articles.
  • Trace monitoring and oversight mechanisms are challenged when internal traces lack legibility.

Proposed Mitigations

The research paper establishes that security issues arise because reasoning block envelopes authenticate content but neglect context. Cross-user replay can be cheaply mitigated without server-side storage by embedding an account identifier directly into the authenticated data at issuance. To secure cross-session contexts without breaking portability features like conversation forking or turn compaction, the authors propose a hash chain paired with a Merkle tree. For existing blocks signed without context, rotating cryptographic keys is the primary retroactive fix, alongside API gateway model version verification and refusal training.

  • Reasoning block envelopes authenticate content but do not authenticate the context in which a block was produced or replayed.
  • Cross-user replay attacks can be prevented without server-side storage by embedding account identifiers into authenticated issuance data.
  • Binding blocks to complete transcripts would break legitimate features like forking conversations, compacting history, and model downgrades.
  • A combination of hash chains and Merkle trees enables session binding and ordering integrity while allowing pruned histories.
  • Addressing previously published blocks requires key rotation, which unavoidably invalidates legitimate continuations of older sessions.
  • Supplemental mitigations include server-side storage, model version verification at API gateways, and training models to refuse transcription requests.

Conclusion

The security issues discussed do not stem from a failure of cryptographic techniques, but rather from a lack of binding between an encrypted block and the context that generated it. A displayed answer summary is distinct from the underlying trace and can diverge from it, while sanitization cannot safely clean uninspectable encrypted blocks without removing them completely. Furthermore, the overall security of a model family is dictated by its least protected model, meaning anti-distillation safeguards on flagship models are undermined if cheaper sibling models accept the same blocks.

  • Cryptographic envelopes functioned properly, but lacked authenticated binding between blocks and their generating context.
  • A summary displayed next to an answer is an independent artifact from the trace it describes, allowing the two to diverge.
  • Session log sanitization only affects plaintext, requiring encrypted blocks to be eliminated entirely since they cannot be inspected.
  • A model family's security depends on its least protected member, reducing the efficacy of flagship anti-distillation training if cheaper sibling models accept identical blocks.