OWASP LLM Top 10 Coverage Gaps for Enterprise Deployments
The framework misses enterprise risks in agentic AI, identity governance, and attack chaining.

The OWASP Top 10 for LLM Applications is the most widely used shared vocabulary for talking about AI security risk, and it earns that status because it is plain-spoken. The project started in 2023 as a community effort, built from practitioner input rather than handed down by a single vendor or regulator. By late 2024, OWASP had updated the list to reflect real incidents, new attack techniques, and the fast growth of agentic AI, systems that do not just answer questions but take actions. The 2026 edition pushes that shift further: the list now spends more attention on what happens after a model is fooled. Prompt Injection still holds the No. 1 spot, but Excessive Agency jumped from sixth to third, a sign that OWASP is watching what a manipulated model is allowed to do as closely as whether it can be manipulated. Unbounded Consumption climbed four places, reflecting how a single request can now set off a chain of costly tool calls. OWASP has been direct about how the list gets built: the public incident record is incomplete, reporting habits differ across organizations, and some categories overlap. That makes the list a prioritized guide for attention, not a precise scorecard of risk.
The List's Design as an Awareness Baseline
OWASP describes the LLM Top 10 as a minimum-awareness document: it names categories of risk without specifying the controls an organization must build in response. The categories on the list describe ways things go wrong. No regulation currently points to this list by name, so nothing in law or in standard audit practice forces anyone to build a specific control because OWASP named a risk. OWASP's 2026 edition does draw a clearer line between an LLM that sits inside an application answering questions and an agent that uses tools, remembers past actions, works with other agents, and causes effects downstream. What it does not do is explain how to govern both of those at once, inside one deployment, with one security team.
How Agentic Deployments Changed the Risk Geometry
The original OWASP list was built around a model that responds: someone asks a question, the model answers, and the risk lives inside that exchange. That is not what most enterprises are running anymore. An LLM application today can pull files, query a database, write and run code, send messages, and kick off business processes on its own. That changes what a successful attack costs a company. A manipulated agent, with real tool access, can reach private data, call an overprivileged tool, or take an action that cannot be undone. Once an agent gains real tool access, prompt injection becomes an action problem: the danger is not what the model says, but what it does with the authority it has been handed. Agentic systems also raise identity questions that older service-to-service security models never had to answer cleanly. Whose authority does the agent act under? Does it carry the identity of the person who triggered it, or does it have its own? Can anyone trace the chain of delegation after the fact? Traditional security models were built for code that behaves the same way every time, where a given input produces a predictable, binary result. Large language models are probabilistic and sensitive to context, so a framework built by looking backward at past incidents will keep lagging behind a technology whose attack surface is still being mapped out in real time.
The EchoLeak case: what a real enterprise attack looks like across OWASP categories
EchoLeak, tracked as CVE-2025-32711 and disclosed in June 2025, shows what happens when an attack does not respect OWASP's category boundaries. The exploit was a zero-click vulnerability in Microsoft 365 Copilot: an attacker could send a single email and steal confidential data without the victim clicking anything. The attack chained three separate OWASP categories together. LLM01, prompt injection, got the malicious instruction into the system. LLM02, sensitive information disclosure, let that instruction pull out confidential data. All three worked in sequence, not as separate incidents but as one connected chain. Microsoft rated the vulnerability 9.3 out of 10 on the CVSS scale, a critical score. The deeper problem appears after the attack, in the forensic record an incident response team must reconstruct. Without logging built specifically for AI decision chains, tool calls, and state changes, an incident response team would struggle to prove the attack happened at all, let alone trace what data left the building. EchoLeak stands as the first known case of a prompt injection weaponized to pull real data out of a production AI system. It sets the bar for what enterprise incident response now has to be able to handle. It also makes a point about the framework itself: each OWASP category involved in EchoLeak was named correctly and described well on its own, but the list reads each one in isolation. Nothing in the framework coordinates a defense across categories once an attacker starts chaining them.
Six specific coverage gaps enterprises must close beyond the OWASP checklist
The OWASP LLM Top 10 names real risks without specifying where the controls for those risks need to live, what they need to check, or how they fit together across an actual production stack. That gap splits into six parts enterprises need to address directly.
The first is agentic identity and privilege governance. Excessive Agency, now LLM03 on the 2026 list, correctly flags the danger of giving a model too much permission. It does not address the identity questions that follow once an agent starts acting: whose authority it carries, whether anyone can audit the chain of delegation, and how its credentials get scoped and rotated. This sits as much in identity and access management and privileged access management as it does in AI security, and a single prompt failure in a system with real tools and credentials can turn into a business-impacting, sometimes irreversible event. The OWASP Agentic AI Top 10, released in late 2025, adds coverage here as a supplement to the original list. Enterprises running multi-agent workflows need both documents at once, stitched together by hand.
The second is audit logging built for AI decisions and tool calls. The LLM Top 10 never mandates a logging architecture. Without logging built for how AI systems actually operate, a company cannot confirm an attack took place, cannot reconstruct what data was touched, and cannot answer an auditor who asks for agent activity logs. EchoLeak's forensic problem is a direct result of that absence, and a large share of organizations today cannot track or audit what data their AI agents are accessing.
The third gap is visibility into vendor and shadow AI. The list addresses supply chain risk at the model and dependency level, but it says nothing about what happens when employees quietly adopt unvetted models, third-party wrappers, and plugins that never pass through security review. Traditional third-party risk management programs were not built to assess AI-specific attack surfaces either, which leaves the vendor AI ecosystem as a structural blind spot without a tool built for the job.
The fourth gap is at the RAG layer, where retrieval-augmented generation pulls in outside documents to ground a model's answers. Sensitive Information Disclosure and Vector and Embedding Weaknesses both touch this part of the stack, but neither entry says where the access control has to live. The permission model in the source system and the permission model inside the AI retrieval system end up running as two separate things.
The fifth gap concerns tool-call authentication, especially in protocols like the Model Context Protocol, or MCP, which connects agents to the tools they use. The control has to move to the moment of the tool call itself, not to inspecting the prompt beforehand. The 2026 list's treatment of Excessive Agency names this risk in general terms but does not specify MCP-layer controls, and no current OWASP document treats tool-call authentication as a first-class requirement. An agent can pass every prompt-level check OWASP describes and still make an unauthenticated call that reaches a production system.
The sixth gap is business-logic abuse that crosses categories. OWASP's entries are written to be read one at a time, but real attacks, as EchoLeak showed, chain several categories together and exploit logic specific to one company's workflow. That kind of abuse sits at the intersection of multiple categories and will not get caught by mapping against any single entry on the list. Catching it takes threat modeling built around the specific application and workflow, something the list does not prescribe.
The Strongest Objection to This Argument
The strongest challenge to this argument is methodological. It holds that heavy investment in OWASP-aligned defenses may already be closing these gaps for serious enterprises, which would mean the framework works even in places it looks incomplete on paper. OWASP reads that gap as a sign of active defense: mature security teams are already spending heavily to prevent and contain prompt injection. That spending is why it appears less often in the incident record than practitioners expect. If that reading holds, OWASP-aligned programs may be doing more real work than the coverage gaps suggest on their own.
That argument holds only for the risks the framework actually covers. Vendor AI ecosystems, shadow AI tools quietly adopted inside a company, RAG-layer access control, and MCP-layer authentication are not places where heavy OWASP-driven investment is happening, because the list never asks for it. The defense effect is real where it applies, but it only protects the ground the map actually covers.
What OWASP's Newer Companion Documents Add
Two OWASP releases from late 2025 into 2026 start to close some of these gaps, but both are additions that enterprises have to actively combine with the original Top 10. The OWASP Top 10 for Agentic AI Applications, released in late 2025, is a separate framework for systems where an LLM plans, decides, and carries out multi-step tasks on its own using outside tools. One entry on the Agentic list, ASI07, Insecure Inter-Agent Communication, gives the risk of spoofed, replayed, or unauthenticated messages between agents its own category for the first time, rather than leaving it as an afterthought buried inside a broader risk.
The second release, the OWASP Agent Control Standard, tries to standardize how a company actually stops or limits an agent once a risk has been identified, built to work across the fragmented landscape of different agent frameworks currently in use. Where the Top 10 tells a security team what can go wrong, the Agent Control Standard addresses how to stop it, filling in some of the control detail that the Top 10's focus on naming failure modes leaves open.
Putting all three documents together, while also mapping the result against the EU AI Act, ISO 42001, SOC 2, DORA, and the NIST AI Risk Management Framework, is a substantial piece of work. It is not the same thing as having one unified control architecture, and most security teams are not staffed to do it by hand without dedicated tooling built for the purpose. An AI Bill of Materials covering agents, tools, and connected servers is becoming the expected artifact as auditors start asking for agent activity logs, permission reviews, and kill-switch procedures, but no OWASP document currently requires one as a deliverable.
Using the OWASP LLM Top 10 as a Starting Point
The enterprises that treat the OWASP LLM Top 10 as a living baseline, one they map explicitly into existing governance frameworks and extend with dedicated AI risk assessment, will close the gaps the list's design leaves open. The ones that treat OWASP alignment as a finish line will not. In practice, that means mapping OWASP categories to specific controls inside frameworks like ISO 42001, the NIST AI RMF, and SOC 2, rather than treating OWASP alignment as a compliance destination in its own right. It means treating the Agentic AI Top 10 and the Agent Control Standard as required reading for any deployment where an LLM has tool access, holds memory across sessions, or coordinates with other agents, not as optional extras for teams with time to spare.
It means building or buying an AI asset inventory, an AIBOM, that covers vendor models, plugins, agents, and connected servers, because nothing can be governed before it is visible. It means putting access control directly at the RAG layer as its own architectural requirement rather than assuming LLM02 implies it, with document-level permission metadata embedded before semantic ranking runs on any pipeline that touches sensitive data. It means moving the authentication checkpoint to the tool call itself in any deployment connected through MCP or a similar protocol, since checking prompts alone will not close that door. And it means building audit logging for AI decisions, tool calls, and state changes as a structural requirement from the start, because without it, confirming an incident happened and answering an auditor's questions are both out of reach.
The sharpest exposure sits with the vendor AI ecosystem, where third-party risk management, information security, privacy, and legal teams all meet the same blind spot. Third-party AI tools bring prompt injection, data exfiltration, and compliance risk that generic governance frameworks and OWASP alignment on their own cannot surface. Continuous monitoring of vendor AI assets, both before they are deployed and for as long as they stay connected, is the capability that closes a visibility gap the OWASP list was never designed to fill. Supply chain and LLM risk management needs continuous inventory, automated testing, and policies that enforce themselves, not a checklist filled out once and filed away, because the attack surface keeps changing as vendor AI tools get updated, extended, and wired into new corners of the enterprise.


