LLM Security Basics: The Full Threat Model
Almost all LLM vulnerabilities originate from the architectural reality that models receive instructions and external data in a single token sequence with no structural boundary between them. While traditional software relies on parameterization to strictly separate executable commands from user data (such as in SQL queries), no such mechanism exists for natural language where instructions and data are both plain text. This allows prompt injection to succeed either directly via chat inputs or indirectly through retrieved external content like emails or documents. Relying solely on filtering mechanisms is insufficient, as evidenced by attacks that bypass dedicated injection classifiers. The OWASP Top 10 for Large Language Model Applications maps critical AI security risks directly onto the stages of an LLM data pipeline, spanning input, retrieval, model processing, tools, and output. Each pipeline stage exposes distinct vulnerabilities, including prompt injection, knowledge base corruption, and excessive agency. Unlike localized pipeline stages, supply chain vulnerabilities underpin the entire architecture and affect all downstream components simultaneously. Mapping attacks to these stages clarifies where untrusted data enters and which permissions are abused, guiding effective defense strategies. Attacks targeting a model's interior—such as model theft, training-data extraction, and training-time poisoning—pose real technical risks but are currently low priority for most standard application developers. Demonstrations show these attacks are bounded, expensive, or promptly mitigated by commercial model providers following disclosure. However, responsibility and risk rise significantly for teams that host open weights, fine-tune on proprietary data, or run their own training infrastructure. For downstream developers, common operational vulnerabilities like over-privileged agents present a far greater threat than interior model attacks.
閱讀原文 ↗目錄
Trust Boundaries
Almost all LLM vulnerabilities originate from the architectural reality that models receive instructions and external data in a single token sequence with no structural boundary between them. While traditional software relies on parameterization to strictly separate executable commands from user data (such as in SQL queries), no such mechanism exists for natural language where instructions and data are both plain text. This allows prompt injection to succeed either directly via chat inputs or indirectly through retrieved external content like emails or documents. Relying solely on filtering mechanisms is insufficient, as evidenced by attacks that bypass dedicated injection classifiers.
- LLM inputs concatenate system prompts, user queries, retrieved data, and tool outputs into a single sequence where any token can influence execution as an instruction.
- Traditional systems prevent injection vulnerabilities (such as SQL injection) using parameterization to separate command from data, but no equivalent exists for natural language.
- Prompt injection occurs via direct inputs (chat interfaces) or indirect content (retrieved documents, emails, web pages, or files).
- Any system feature processing retrieved external text inherently carries indirect prompt injection risk by construction.
- The EchoLeak attack demonstrated that indirect prompt injection payloads can bypass dedicated security filters, including Microsoft's cross-prompt-injection classifier.
Attack Surface
The OWASP Top 10 for Large Language Model Applications maps critical AI security risks directly onto the stages of an LLM data pipeline, spanning input, retrieval, model processing, tools, and output. Each pipeline stage exposes distinct vulnerabilities, including prompt injection, knowledge base corruption, and excessive agency. Unlike localized pipeline stages, supply chain vulnerabilities underpin the entire architecture and affect all downstream components simultaneously. Mapping attacks to these stages clarifies where untrusted data enters and which permissions are abused, guiding effective defense strategies.
- The OWASP Top 10 for LLM Applications maps specific security risks across sequential pipeline stages: Input, Retrieval, Model, Tools, Output, and underlying Supply Chain.
- Input-stage risks include direct prompt injection and unbounded consumption, also termed 'denial of wallet.'
- The 2024 study PoisonedRAG achieved a 90 percent attack success rate on targeted questions by injecting only five malicious passages into a knowledge base containing millions of documents.
- Excessive agency occurs at the tool stage when an agent is granted more permissions than necessary to complete its task.
- Supply chain threats span the entire pipeline simultaneously, as a compromised model or vector store impacts every subsequent stage.
Model Attacks
Attacks targeting a model's interior—such as model theft, training-data extraction, and training-time poisoning—pose real technical risks but are currently low priority for most standard application developers. Demonstrations show these attacks are bounded, expensive, or promptly mitigated by commercial model providers following disclosure. However, responsibility and risk rise significantly for teams that host open weights, fine-tune on proprietary data, or run their own training infrastructure. For downstream developers, common operational vulnerabilities like over-privileged agents present a far greater threat than interior model attacks.
- Recovering a production model's final embedding layer cost under twenty dollars, but reconstructing a complete frontier model via API remains cost-prohibitive compared to training one.
- In late 2023, researchers demonstrated training data extraction from ChatGPT by prompting it to repeat a single word continuously, prompting OpenAI to filter the vulnerability.
- A 2025 study showed roughly 250 malicious documents could backdoor models from 600 million to 13 billion parameters, disproving the assumption that larger models need proportionally more poison data.
- Model interior attacks pose greater operational danger to teams managing their own weights, fine-tuning pipelines, or self-hosted models than to API consumers.
- Organizations risk security misallocations when prioritizing rare interior attacks over high-frequency operational risks such as over-permissioned agents.
Excessive Agency
LLM attacks that cause material damage exploit a structural vulnerability known as the lethal trifecta, which occurs when an agent possesses access to private data, exposure to untrusted content, and a channel to act externally. Model alignment cannot prevent these exploits because LLMs are inherently designed to follow instruction-like inputs. Real-world incidents have already compromised systems like GitHub's and Anthropic's MCP implementations, GitLab Duo, and autonomous trading agents. Mitigating this risk is typically achieved most affordably by narrowing access permissions or severing outbound communication channels rather than relying on stronger filtering.
- The lethal trifecta consists of three co-occurring capabilities: private data access, untrusted content exposure, and an external action or exfiltration channel.
- Model alignment does not eliminate vulnerability to injected instructions because models inherently follow instruction-like prompts.
- GitHub's Model Context Protocol (MCP) server was exploited via malicious public repository issues to leak private repository data.
- GitLab Duo leaked private repository contents after being fed a public project containing hidden instructions.
- Anthropic's official Git MCP server was assigned three injection-related CVEs in 2025.
- Eliminating an agent's outbound channel or narrowing data access is typically cheaper and more effective than implementing stronger input filters.
Supply Chain
Supply chain vulnerabilities in AI stacks pose severe risks by bypassing runtime defenses prior to input validation. Many distributed models use serialization formats like Python pickle that allow arbitrary code execution upon loading, exemplified by the nullifAI technique evading Hugging Face's Picklescan. The widespread scope of this threat is demonstrated by Protect AI flagging roughly 352,000 issues across over 50,000 models. Mitigating this risk requires strict provenance management, adoption of safe serialization formats, and model signing.
- Supply chain compromises bypass runtime defenses because malicious code executes when serialized files load, before input validation runs.
- In early 2025, ReversingLabs documented nullifAI, a technique embedding reverse shells in compressed Python pickle files to evade Hugging Face's Picklescan scanner.
- Protect AI inspected over four million models hosted on Hugging Face and flagged roughly 352,000 suspicious or unsafe issues spanning more than 50,000 models.
- Safer serialization formats that do not execute arbitrary code upon loading serve as a critical defense against model supply chain attacks.
- Model signing verifies origin and integrity, mirroring established secure release practices in software package ecosystems.
Defense in Depth
A single security guardrail cannot reliably defend AI models against adaptive attacks, necessitating a defense-in-depth architecture where multiple independent layers prevent single-point failures. Empirical research from late 2025 demonstrates that adaptive attacks can defeat multiple existing prompt injection and jailbreak defenses. Rather than attempting to make the model itself impenetrable, effective security architectures isolate and constrain the model's environment, treating its outputs as untrusted. Practical implementations of this philosophy include frameworks like Google DeepMind's CaMeL and operational heuristics like Meta's Agents Rule of Two.
- A joint November 2025 study by OpenAI, Anthropic, and Google DeepMind showed that adaptive attacks defeated twelve existing defenses against prompt injection and jailbreaking.
- Lone guardrails and production filters cannot guarantee complete protection, shifting the realistic operational goal from total prevention to breach survivability.
- Durable AI security involves constraining the surrounding system and treating the model as untrusted rather than relying purely on model self-defense.
- Google DeepMind's CaMeL architecture uses a privileged planner and quarantines external data to block unverified actions.
- Meta's Agents Rule of Two limits an autonomous agent to satisfying no more than two of three risks: untrusted input, sensitive access, and unreviewed external actions.
- Effective multi-layered defenses incorporate input validation, data quarantine, least-privilege tool access, output sanitization, monitoring, and human-in-the-loop review.
Conclusion
The foundational threat model for language models arises because they process instructions and data as an indistinguishable sequence of tokens. Attacks like model theft and training-data extraction are largely bounded, but catastrophic risk concentrates when an AI agent simultaneously accesses private data, untrusted content, and external communication channels. Mitigating these risks demands a defense-in-depth approach across the supply chain, though safeguards such as human review inevitably trade off autonomy, latency, and cost.
- The fundamental root cause of LLM security vulnerabilities is the absence of an architectural boundary between instructions and data.
- Major operational risk is concentrated where an agent simultaneously handles private data, untrusted content, and external channels.
- Interior attacks, including model theft and training-data extraction, are largely bounded and mitigated.
- Supply chain provenance represents the attack surface most directly controllable by system operators.
- Defense in depth is essential because no single protective layer holds, though human review inherently reduces operational autonomy.