Est.

Prompt Injection Defense Gaps in OWASP LLM Top 10

OWASP's defenses miss indirect, agentic, and adaptive injection attacks.

Senior Writer · · 12 min read
Cover illustration for “Prompt Injection Defense Gaps in OWASP LLM Top 10”
Prompt Injection · September 29, 2026 · 12 min read · 2,702 words

The OWASP LLM Top 10 correctly identifies prompt injection as the top threat, but its prescribed defenses were built for direct, input-level attacks and leave meaningful gaps when attacks are indirect, agentic, or adaptive, gaps that enterprises relying on the framework as a compliance checklist are likely to miss.

Why prompt injection sits at the top of OWASP's LLM risk list

OWASP's LLM Top 10 was first released in 2023, updated for 2025, and a 2026 edition has since also been released. That newest edition still keeps ten broad risk categories, but it reorders them and for the first time folds actual incident data into how the ranking gets built. Prompt injection has the worst track record of any risk on the list. It's the only risk in the entire list that has never moved down, and there's a reason for that.

The reason comes down to how language models actually process what's in front of them. A model reads instructions and data as one continuous stream of tokens, with nothing in the architecture flagging where a system command ends and a user's document begins. That's not a bug some patch will fix next quarter. It's a structural condition of how these systems work. OWASP's own definition gets at this directly: a prompt injection vulnerability happens when user prompts alter the model's behavior or output in ways nobody intended, and the injected text doesn't even need to be readable by a human, only parseable by the machine. That second part matters more than it sounds. An attacker doesn't need clever phrasing. They need something the model will tokenize and act on, which is a much lower bar.

So what does it mean that this risk has held the top spot through two full revisions? Consensus, mostly⟧c8⟧. Security researchers, vendors, and enterprises broadly agree this is the most urgent problem in the category. But agreeing on the diagnosis isn't the same as agreeing on the cure, and that gap is where this piece spends its time. A successful injection can leak sensitive data, hand an attacker unauthorized access to whatever systems the model touches, warp a decision the model was trusted to make, or push out harmful content dressed up as legitimate output. OWASP gets the threat right. The open question, the one to chase through the rest of this, is whether the defenses it prescribes actually cover the ground enterprises are now operating on.

What OWASP's framework prescribes as defenses and quietly concedes won't fully work

OWASP splits prompt injection into two flavors. Direct injection is the simple case: someone types malicious instructions straight into the prompt box, and the attacker controls that input from the start. Indirect injection is stranger and harder to police, because the attacker never touches the user's keyboard at all. Instead they plant instructions inside a document, a web page, an email, or some entry in a knowledge base, and the model picks up those instructions later when it retrieves that content on someone else's behalf. The person using the assistant did nothing wrong. They just asked a question, and the answer came back poisoned.

Constrain the model's behavior through a well-built system prompt. Define the output format ahead of time and check it with deterministic code, not vibes. Run input and output filtering, including semantic filters, string checks, and something called the RAG Triad, which scores a response on context relevance, groundedness, and how well the answer actually matches the question. Keep the model on a short leash with least-privilege access, so it never holds credentials it doesn't strictly need. Route anything high-risk or irreversible through a human for approval. Label external content clearly so untrusted text never gets mistaken for a trusted instruction. And test the whole system adversarially, on a recurring basis, treating the model itself like an untrusted user rather than a partner.

That's a reasonable list. It's also, by OWASP's own admission, incomplete. Research shows neither RAG nor fine-tuning fully mitigates prompt injection vulnerabilities, which is a fairly striking thing to see printed inside the document that's supposed to be the fix. OWASP goes further still, conceding it's unclear whether any fool-proof prevention method even exists for this class of vulnerability. Defense in depth is the strategy on offer, not a guarantee.

Most of these controls were built around the direct, single-turn version of the attack, the one where a human types something bad into a box. Practitioners aligned with OWASP's guidance lean on a familiar detection layer, including things like Azure Prompt Shields for real-time screening, Llama Guard 3 for classifying whether an input or response is safe, and LLM-Guard style scanners sitting on both ends of the pipeline. These tools do real work. But they're a layer, not a wall, and the next section gets into why the wall keeps developing new doors. OWASP's prescribed defenses (per LLM01:2025) are laid out, alongside what the framework quietly concedes won't fully work.

How the attack surface shifted when LLMs became agents with tools

Somewhere in the last couple of years, language models stopped just answering questions. They started sending emails, querying databases, calling APIs, browsing live web pages, executing code, editing files, and kicking off entire workflows on their own. That shift changes what a successful attack actually costs a victim.

When a model is only a chatbot, a successful injection produces a bad answer. Maybe embarrassing, maybe a data leak, but the damage stays roughly where it started. When that same model has tool access, a successful injection hands the attacker the tools themselves. The blast radius jumps by an order of magnitude, because now the attacker isn't just manipulating text, they're manipulating actions, and actions touch real systems.

OWASP's 2026 edition tracks this shift, even if its core LLM01 mitigations haven't fully caught up to it. Excessive Agency climbed from sixth place to third, which is the framework quietly admitting that giving models too much autonomy is now a top-tier concern in its own right, not a footnote under prompt injection. And the deployment data backs this up: roughly 53% of companies skip fine-tuning entirely and lean on retrieval-augmented generation and agentic pipelines instead, according to figures from confident-ai.com OWASP Top 10 for LLM Applications (2025). That means the indirect-injection surface, the one built on retrieved content rather than typed input, is the dominant pattern in production today, not some edge case worth a footnote OWASP Top 10 for LLM Applications (2025).

Researcher Simon Willison gave this problem a name in 2025 that stuck: the lethal trifecta. It has access to private data. It gets exposed to untrusted external content. If any one of those three legs is removed, the attack path collapses on its own. If all three remain standing, the agent is exploitable no matter how good the input filters sitting in front of it look on paper.

That's the uncomfortable part. Filters catch bad text. They don't catch an agent that was never supposed to have all three of those properties in the first place. RAG systems, browsing agents, MCP tool servers, and email-summarizing assistants all consume attacker-controllable text (indirect injection is now the center of the threat model, not an edge case). Indirect injection is the center of the threat model. It's the center of the threat model, and retrieval plus tool-use has simply outpaced the defenses meant to hold them in check.

Diagram: The Lethal Trifecta: When All Three Legs Stand. Visualizes: Visualize Simon Willison's 'lethal trifecta' concept: an LLM agent becomes exploitable when three conditions coexist simultaneously — (1) it has access to private data, (2) it is…

Production incidents that moved prompt injection from theoretical to actively exploited

Before mid-2025, most of this stayed academic. Researchers demonstrated attacks in papers and conference talks, but nobody could point to a confirmed breach at real scale. That changed fast, and it changed against products built by companies with serious security budgets.

It stands as the first confirmed zero-click data exfiltration vulnerability against Microsoft 365 Copilot. The attack vector was almost insultingly ordinary: a crafted email, no attachment, no link for anyone to click. There was nothing for employee training to catch, because the employee didn't do anything. Copilot simply pulled data from OneDrive, SharePoint, and Teams during normal operation and exfiltrated it by routing through a trusted Microsoft domain, using SharePoint and Teams URLs as a relay to an attacker-controlled server. The attack slid past Microsoft's own cross-prompt-injection classifiers and its link redaction system entirely, because the egress traffic passed every destination check. The domain it left through was still sitting on the allowlist. This is the case that illustrates precisely the gap OWASP's input-filtering guidance doesn't cover: the injection arrived in a channel (email body) that no filter was watching, and the egress path was permitted by design.

GitHub Copilot had its own incident in early 2025. Researchers found that opening a malicious file in VS Code was enough to inject instructions the assistant would then carry out, resulting in remote code execution on the developer's own machine.

Cursor's IDE saw something structurally similar. An indirect prompt injection exploited a missing confirmation step for new workspace settings files, and Cursor's own AI agent created a malicious .cursor/mcp.json configuration file without asking the user first.

Then there's the Unit 42 finding from Palo Alto Networks, first detected in December 2025 and published in early 2026. It documents the first observed case of malicious indirect prompt injection used in the wild specifically to bypass an AI-based ad review system, a narrower and more concrete claim than the vague "attacks are happening somewhere" framing that circulated before it.

The CVSS pattern across incidents (9.3, 9.6, 9.8) is not incidental vectra.ai. These are critical-severity vulnerabilities in enterprise-grade, heavily resourced vendor products that had already implemented OWASP-aligned controls vectra.ai. None of these incidents were caught by the input validation, output filtering, or content segregation controls OWASP recommends. They exploited the seams between those controls and the surrounding system. Each one found the seam between two controls, the place where one system's assumption stopped and another's began, and walked straight through it. The incident was disclosed mid-2025 by Aim Labs, and Microsoft assigned CVE-2025-32711. The CVSS score was 9.3, according to vectra.ai. The CVSS score was 9.6, according to vectra.ai. The attacker achieved remote code execution through the malicious MCP configuration. The CVSS score was 8.6, according to vectra.ai.

Diagram: Real-World Attack Success Rates Against Deployed Defenses. Visualizes: Visualize the published attack success rates against specific deployed defenses to show how brittle current protections are under adversarial pressure.

Start with a number that should stop most compliance conversations cold. Against systems that already have defenses deployed, prompt injection attacks succeed somewhere between 50% and 84% of the time, depending on configuration and how many attempts the attacker gets, per figures from Vectra AI vectra.ai. That's not a system with no protection. That's the baseline with protection running.

The first structural gap is pattern matching itself. SQL injection has a schema to check against. Natural language doesn't, and that flexibility is precisely what makes it hard to police with a filter built to spot fixed patterns. Researchers demonstrated evasion rates as high as 100% against prominent protection systems, including Microsoft's own Azure Prompt Shield and Meta's Prompt Guard, according to research compiled by introl.com AttriGuard: Defeating Indirect Prompt Injection in LLM Agents via Causal Attribution of Tool Invocations. A filter that a determined attacker can defeat every time isn't really a filter anymore; it's a speed bump.

The second gap appears specifically against adaptive attackers, the ones who adjust their approach based on what a defense blocks. A 2025 NAACL Findings paper by Zhan, Fang, Panchal, and Kang tested eight separate defenses against indirect prompt injection in LLM agents, and bypassed every single one using adaptive attacks, holding success rates above 50% consistently vectra.ai aclanthology.org AttriGuard: Defeating Indirect Prompt Injection in LLM Agents via Causal Attribution of Tool Invocations. A follow-up piece of research, Nasr et al., pushed the same logic into jailbreak defenses and reported attack success rates above 90% across twelve published defense mechanisms arxiv.org sysdig.com AttriGuard: Defeating Indirect Prompt Injection in LLM Agents via Causal Attribution of Tool Invocations. A detector that scores better on a fixed benchmark isn't measurably harder for an adaptive attacker to get around than the version before it. Static testing keeps producing prettier benchmark numbers while the real-world resistance barely moves.

Fine-tuning runs into its own ceiling. SecAlign, one of the stronger published fine-tuning defenses, still misses roughly one in ten optimization-based attacks. A newer approach called ReasAlign, released in January 2026, cut attack success down to 3.6% on one open-ended benchmark, but that number comes from a static test, not from an attacker adapting specifically to beat ReasAlign sysdig.com. Before an approach called SecOPD arrived, the prior state of the art, Meta-SecAlign, was getting beaten 94% of the time by adaptive PISmith attacks SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation AttriGuard: Defeating Indirect Prompt Injection in LLM Agents via Causal Attribution of Tool Invocations. The leading fine-tuned defense in the field was failing on almost every serious attempt. SecOPD, built at UC Berkeley using token-level feedback distillation, brought that down to 9.0% against PISmith and 4.7% in agentic tool-calling scenarios, a real improvement, though those numbers still come from fixed benchmarks SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation AttriGuard: Defeating Indirect Prompt Injection in LLM Agents via Causal Attribution of Tool Invocations.

Multimodal inputs open a fourth gap that most text-focused filters simply can't see. Yeo and Choi demonstrated that image-based prompt injections walk right past filters built to scan text vectra.ai. Any domain where a model reads non-textual data, medical imaging being the obvious example, carries exposure that current OWASP guidance doesn't directly cover.

And then there's the RAG-specific surface. Embedding models themselves can get manipulated to return misleading similarity results, which quietly undermines the trustworthiness of everything a retrieval system produces downstream. OWASP added Vector and Embedding Weaknesses as its own category, LLM08, in the 2025 edition. But the LLM01 guidance on prompt injection hasn't caught up to treat these as attack paths that compound with injection rather than sit next to it. A second gap is that published defenses are brittle against adaptive attackers. Gap 3 is that fine-tuning has a hard ceiling against optimization-based attacks. Gap 4 is that multimodal inputs bypass text-only filters entirely.

Why enterprises using OWASP as a compliance checklist face the widest exposure

The uncomfortable part for anyone running a security program is this. OWASP's own format, numbered items, prescribed mitigations, a defensible audit trail, is exactly the shape a compliance team wants to see. It's built for checking boxes. But satisfying a control and actually closing the attack surface behind that control are two different outcomes, and the incidents above prove it. EchoLeak breached an enterprise that had Microsoft's own classifiers running the whole time. The Cursor attack didn't slip past a filter at all. It walked through a missing confirmation step that nobody had flagged as a gap.

OWASP does recommend adversarial testing, to its credit. Most enterprise third-party risk management and infosec programs assess vendor AI at a single point in time, a snapshot. They don't continuously watch for new attack surface introduced every time a model gets updated, a new tool integration gets bolted on, or a new MCP connector goes live.

That gap widens further once you account for how most enterprises actually consume this technology. They aren't only building their own models. They're running vendor AI, Copilot, Cursor, Slack AI, assorted coding assistants, where they have no visibility into the model's internal controls, its retrieval pipeline, or what permissions its tools actually carry. A standard TPRM questionnaire asks a vendor whether security controls exist. It doesn't ask, and typically can't ask, whether those controls hold up against an adaptive indirect injection routed through that vendor's own RAG pipeline. That's a blind spot generic governance paperwork was never built to catch.

The lethal trifecta gives a sharper diagnostic than any checklist item OWASP publishes: does this agent hold private data, get exposed to untrusted content, and retain the ability to communicate outward, all three at once? Answering that honestly takes continuous visibility into how an agent is actually configured and what it can reach. OWASP got the threat right by putting prompt injection at the top of the list. Whether an organization's defenses actually match that threat is a separate question, and it's one no framework can answer on a company's behalf.

Sources

  1. LLM01:2025 Prompt Injection
  2. SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation
  3. AttriGuard: Defeating Indirect Prompt Injection in LLM Agents via Causal Attribution of Tool Invocations
  4. vectra.ai
  5. sysdig.com
Filed underPrompt Injection

More in Prompt Injection