Est.

Insecure Output Handling Vulnerabilities in LLM-Integrated Applications

Treat all LLM-generated output as untrusted input before it reaches downstream systems.

Contributing Writer · · 10 min read
Cover illustration for “Insecure Output Handling Vulnerabilities in LLM-Integrated Applications”
LLM Vulnerabilities · October 11, 2026 · 10 min read · 2,207 words

It is a structural condition built into every LLM-integrated application: the model produces a string, and that string must be treated as untrusted input the moment it leaves the model, no matter how polished the response looks.

Why LLM output cannot be treated as trusted data

Picture the actual mechanics for a second. A model generates text, and that text flows somewhere: a browser renders it, a shell executes it, a database driver runs it as a query. The security of the entire feature rests on what the application does with that string in the instant before it reaches its destination. OWASP's LLM05:2025, Improper Output Handling, names this directly: insufficient validation, sanitization, and handling of LLM outputs before they get passed to downstream components and systems.

Inside the model, there's no wall separating the system prompt from the user's input, from a retrieved document, from a tool's output, from the conversation history. Next-token prediction doesn't know or care which part came from an administrator and which part came from a stranger's email. A language model enforces nothing of the sort internally, so the application has to build that boundary itself, at the point where it consumes whatever the model hands back.

This is a different problem from Overreliance, which is about broader overdependence on whether an LLM's answers are accurate or appropriate. The vulnerability doesn't live inside the model's weights. It lives in application code, in the decision (made by a developer, often implicitly) about whether a generated string gets to reach a renderer, an interpreter, or a database driver without a checkpoint in between.

Attack classes mapped to downstream consumers

What determines the kind of attack that's possible? The consumer. Wherever unvalidated output lands, that destination shapes the exploit. A string handed to a shell produces a different failure than the same string handed to a browser. Every consumer boundary deserves to be reasoned about on its own terms rather than lumped into one generic "output handling" bucket.

OWASP's LLM05:2025 lays out five canonical attack classes, each tied to a specific type of consumer. Output fed directly into a system shell or an exec/eval function can produce remote code execution. And content generated for email templates, inserted without proper escaping, can produce XSS against recipients using vulnerable email clients.

The RCE class has a documented real-world instance: in Vanna.AI's CVE-2024-5565, the model wrote SQL, the results were passed to Plotly for visualization, and the application ran exec() on Plotly code the model had generated. A prompt injection at that point didn't just manipulate a chart; it could produce arbitrary Python execution, because the trust boundary between "generated visualization code" and "code safe to execute" simply didn't exist.

Markdown image exfiltration is a subtler variant. A string like an image tag pointing to an attacker-controlled URL with query parameters attached can leak data the instant a client auto-fetches that image, so no script tag is required. That's the kind of attack that slips past defenses built only to catch <script> blocks.

Successful exploitation across the OWASP scenario spectrum

None of this is theoretical once it succeeds. OWASP confirms the consequence spectrum directly: exploitation can produce XSS and CSRF in web browsers, and SSRF, privilege escalation, or remote code execution on backend systems. The severity in any given case tracks the privileges the LLM was granted to begin with.

OWASP's documented scenarios make the range concrete. In one, a general-purpose LLM passes its response to a browser extension without output validation, and the extension shuts itself down for maintenance, a case of privilege abuse through a trusted consumer. In a fourth, a web app renders LLM-generated content with no sanitization, an attacker submits a crafted prompt, and the resulting JavaScript payload runs in every victim's browser. In a fifth, someone manipulates an LLM generating dynamic email templates into embedding malicious JavaScript, and it fires in vulnerable email clients. And in a sixth, a software company lets an LLM generate code from natural language, a workflow that carries risk of SQL injection, insecure data handling, and hallucinated packages that developers might actually install.

What ties these together? The privilege the application hands the LLM. OWASP is explicit that when an application grants a model more privilege than the end user it's acting on behalf of would normally have, an output-handling bug becomes a path to privilege escalation. That variable, how much the LLM is trusted to do on its own, is the one to watch, because one 2025 incident showed what happens when it's set too high in a system used by millions.

EchoLeak proved this wasn't a hypothetical risk confined to OWASP's scenario list. A single crafted email, sent with no interaction required from the victim, was sufficient to exfiltrate data from a Copilot user's entire accessible context: Outlook, Teams, OneDrive, SharePoint, all of it.

How does an attack reach that far without a click? The mechanism, documented in the Aim Security paper (arXiv:2509.10540), was a compound bypass chain, not a single flaw. First, the attack evaded Microsoft's XPIA (Cross Prompt Injection Attempt) classifier, the system built specifically to catch this category of attack. Second, it got around link redaction by using reference-style Markdown instead of the inline format the redaction logic expected. Third, images referenced in content get auto-fetched, and it exploited that. Fourth, it abused a Microsoft Teams proxy that the application's own Content Security Policy permitted to run, turning a sanctioned channel into the exfiltration path.

Once that chain completed, everything Copilot could reach was in scope. Sentra has described EchoLeak as the first documented case of prompt injection weaponized for concrete data exfiltration in a production AI system, and the exposure wasn't unique to Copilot's codebase. Any LLM-based assistant with access to multiple internal data sources has this as a structural feature. Microsoft rated the vulnerability CVSS 9.3, critical, and patched it in June 2025, with no evidence found of exploitation in the wild before disclosure.

Indirect prompt injection, instructions smuggled into content that an assistant retrieves and processes as part of its normal job, is how retrieval-based AI assistants work, and that lesson outlasts the patch. It did not, and could not, eliminate the underlying attack class, because that class is a property of the architecture, not a bug in one implementation of it.

Agentic architectures and lateral movement from a single output-handling failure

Now extend the picture from one assistant to several working together. In a multi-agent system, one agent's output doesn't just get displayed to a user or written to a database. An unvalidated output at agent one is an unvalidated instruction stream handed directly to agent two, and the failure can travel across an enterprise without a single server being breached in the traditional sense.

The Cloud Security Alliance's "Living Off the Agent" research gives this pattern a name: LOTA. The CSA's tracking shows this isn't a fringe concern: lateral movement appeared in none of the documented multi-stage agentic AI incidents in 2023, climbed to a minority share in 2024, and occurred in eight of 21 incidents tracked across 2025 and 2026, tracking the rapid rate at which enterprises are standing up agentic deployments.

What makes this possible structurally? One agent has no reliable way to confirm that the instruction arriving from another agent is legitimate rather than injected, so a single manipulated agent can push adversarial instructions downstream and have them treated as trusted operational directives. The ServiceNow Now Assist case, disclosed by AppOmni in November 2025, shows this exact pattern at the platform level: a privilege escalation path running through the vendor's own agent architecture, one that the organizations using the platform did not design and have no unilateral way to patch.

The Model Context Protocol boundary deserves specific attention here, because this is the point where an output string gets converted into a tool call. Validation belongs at that exact junction, not somewhere upstream of it. Multi-model pipelines create a parallel version of the same unchecked boundary: one model's response becomes another model's context, and whatever got missed at the first output boundary is now sitting inside the second model's working memory as though it were trustworthy.

Diagram: Lateral Movement in Agentic AI Incidents: 2023–2026. Visualizes: Show the rise of lateral movement as a share of documented multi-stage agentic AI incidents across three periods: 2023 (none of the documented incidents involved lateral…

The standard prescription, treat LLM output as untrusted, validate everything, is necessary but not sufficient

If developers follow established appsec practice, output encoding aligned with OWASP's ASVS, parameterized queries, strict CSPs, context-aware sanitization, hasn't the problem already been solved by tools the industry has used for decades?

For a large share of cases, yes. Single-turn, browser-facing LLM features respond well to traditional controls. XSS, SQL injection, path traversal, these are addressable with the same tooling that's protected applications against untrusted strings for years, LLM-generated or not.

But indirect prompt injection breaks that confidence. If instructions get embedded in content the model retrieves and processes, sanitizing the output can't catch them, because the injection has already done its work before the output was ever produced. By the time the application inspects the string coming out of the model, the instruction inside it has already been followed. Sanitizing after the fact catches the wrong moment.

That's the structural lesson EchoLeak left behind: indirect prompt injection is a property of how retrieval-based assistants are built, and patching one CVE doesn't retire the attack class that produced it. The problem compounds further in multi-agent systems, where there's no consistent standard for verifying message integrity between agents, so clean validation at the output boundary of one agent says nothing about the integrity of what the next agent receives. Compounding this further still is the basic unpredictability of the underlying system: because LLM outputs are probabilistic rather than deterministic, a defense that works correctly nearly every time can still fail on a specific input, a different kind of security problem than the deterministic systems appsec was originally built around.

None of this argues for abandoning input validation. It argues for recognizing that validation is a necessary layer, not a complete answer, and that the rest of the answer has to come from somewhere else: defense in depth, layered so that no single control, no matter how well implemented, is left holding the entire weight of the system's security.

Layered controls that reduce the exploitable surface at each consumer boundary

Closing this gap takes controls operating at three separate layers: how the application interacts with the model, what happens to the output string before it reaches a consumer, and the broader architecture that limits damage when the first two layers fail.

At the model interaction layer, OWASP recommends a zero-trust posture: treat the model the same way an application would treat any other user, applying real input validation to whatever comes back before it reaches backend functions. In RAG pipelines specifically, you need to clearly mark external content as untrusted retrieved material, distinct from instructions, so no one can mistake it for an authoritative command.

At the application processing layer, output encoding has to match the destination: HTML encoding for anything rendered in a browser, SQL escaping for anything touching a database, shell escaping for anything near a command line. For MCP and plugin boundaries specifically, you need to validate at the exact point where an output string turns into a tool call, not somewhere further upstream where it's easier to implement but too early to catch the actual risk.

At the system architecture layer, the principle of least privilege does the heaviest lifting: an LLM should never hold more permission than the end user it's acting for, which caps how far an output-handling bug can escalate even when it succeeds. Rate limiting and anomaly detection do the same job from a different angle. Third-party extensions and plugins need independent validation of their own, because OWASP flags inadequate extension input validation as a condition that amplifies everything else on this list.

Continuous monitoring of vendor AI assets is the control generic frameworks miss

Put all three layers in place inside an application; a gap still remains for any enterprise running vendor-supplied AI: copilots, RAG assistants, agentic platforms built and operated by someone else. In that context, you can't fix Improper Output Handling once and move on. It's an ongoing operational risk, because the organization using the system can't see or control everything happening at every downstream consumer boundary inside someone else's product.

EchoLeak is the clearest illustration available. The vulnerability lived inside Microsoft 365 Copilot, a vendor-supplied system. The ServiceNow Now Assist case follows the identical pattern at the platform level: a privilege escalation path running through vendor agent architecture that the organizations deploying it did not design and cannot patch on their own.

The scale of the problem is still expanding. CSA's LOTA research ties the growth in lateral movement incidents directly to the pace at which enterprises are adding agents, connectors, and data sources, and that threat surface doesn't hold still long enough for a one-time review to capture it accurately.

A point-in-time assessment, the kind done at procurement, captures the state of a system at a single moment and nothing after it. The right question for any organization running vendor AI is whether that asset still behaves inside its expected trust boundaries today, and whether anyone would notice the moment it stopped, not whether it passed review at the time it was procured.

Sources

  1. LLM05:2025 Improper Output Handling - OWASP Gen AI Security Project
  2. LLM Insecure Output Handling: An Introduction

More in LLM Vulnerabilities