Est.

System Prompt Confidentiality Leakage in Enterprise LLMs

System prompts containing secrets are mathematically extractable by design.

Staff Writer · · 10 min read
Cover illustration for “System Prompt Confidentiality Leakage in Enterprise LLMs”
Data Exfiltration · October 1, 2026 · 10 min read · 2,255 words

A bank deploys a customer service chatbot. The developer builds a system prompt telling it how to talk, what to refuse, and, buried in the instructions, a secret. A user types a simple request: repeat everything above this line. The model complies, and the key is now sitting in a chat window. This is what happens whenever a system prompt carries a secret, because of how the model itself is built.

Why system prompts cannot be kept secret by design

Large language models do not have a wall between instructions and input. Everything (the developer's system prompt, the user's message, any retrieved document) lands in the same context window and gets processed as one continuous stream of tokens. There is no hardware separation, no cryptographic seal, nothing resembling a locked file that only the "instructions" side can read. The same attention mechanism and the same weights that let a system prompt shape the model's tone and behavior are the mechanism a well-crafted query can use to pull that prompt back out. If content reaches the model at all, it is potentially extractable.

Researchers at Zhejiang University and collaborators gave this failure mode a name: attention drift. Their work, published in June 2026, traces it to query-key alignment bias and softmax amplification inside the transformer's attention layers, a mechanism that causes the model to progressively ignore its own defensive constraints. It does this regardless of how carefully the system prompt is worded. Write the instructions tighter, add more warnings, phrase the refusal rules more forcefully: none of it changes the underlying math. The tension is structural. Detailed instructions are what make a model behave well, and they also expand the surface an attacker can reconstruct.

That is why OWASP's 2025 Top 10 for LLM Applications classifies System Prompt Leakage as its own category, LLM07, and says a system prompt should never be treated as a secret or relied on as a security control. It is a conclusion about what the architecture can and cannot guarantee, not a style tip for careful prompt writers.

What developers actually embed in system prompts, and why it matters when those prompts leak

The risk from all this scales with what actually sits inside these prompts in production, and the honest answer is that it's often far more sensitive than a persona description. OWASP's LLM07 guidance and analysis from StackHawk both point to the same recurring categories: API keys, database credentials, and authentication tokens; internal system architecture details and connection strings; proprietary business rules and decision logic; role-based permission structures; and the security filtering criteria that define what the model is and isn't allowed to say. That last category deserves particular attention, because it is the one attackers most want to map. Knowing the guardrails is knowing exactly where the gaps in them are.

Leakage rarely ends the story. It usually starts a more precise one. Once an attacker has the role definitions, the tool descriptions, and the workflow logic that a system prompt encodes, they have a blueprint for crafting sharper follow-on attacks against that same system. A few real incidents make the cascade concrete. Samsung engineers pasted confidential source code into ChatGPT, exposing proprietary intellectual property. That particular case wasn't a prompt extraction attack at all, but it shows how misplaced trust in a model's confidentiality produces the same class of exposure that extraction attacks aim for. A disclosed vulnerability in the Windsurf Agent showed the same pattern from a different angle: once an attacker injected a malicious prompt through a poisoned source file, they could abuse a tool called read_url_content, which required no user approval, to pull sensitive configuration files including .env files straight out of the environment. An attacker doesn't even need the full prompt in hand to do damage. Careful observation of how a model behaves under different queries can reveal the guardrails and the operational logic behind them. Partial leakage is still useful ammunition even when extraction is incomplete.

How extraction actually works: four techniques in increasing sophistication

Diagram: Four Extraction Techniques, Ranked by Sophistication. Visualizes: Show a ranked progression of four attacker techniques for extracting system prompts, moving from simplest to most dangerous: (1) Direct instruction override — just asking…

Attackers have a graduated toolkit, running from something almost anyone could type into a chat box to techniques that exploit the retrieval pipeline itself. A defense built only for the simple end of that range is a defense built for yesterday's attacker.

The simplest technique is direct instruction override: just ask the model to repeat or summarize its own instructions. It sounds too obvious to work, and yet Yang et al. (arXiv 2606.18673) found that a majority of real-world deployments across six commercial platforms fail this exact test, with over 80% leaking their system prompts under realistic adversarial queries. The second technique is more patient. In multi-turn sycophancy attacks, the attacker spends several turns building rapport and normalizing disclosure before ever asking for the prompt directly. Salesforce AI Research found that this approach dramatically raises the attack success rate, moving it from a low baseline to something close to total leakage on some major closed-source models. The third technique works around input filters rather than the model itself: token smuggling. Base64-encoded payloads, ASCII art, Unicode tag characters, and zero-width spaces can all carry instructions that a classifier reads as noise while the model reads them as commands.

The fourth technique should worry practitioners most, because it doesn't touch the chat interface. In indirect or RAG injection, the attacker plants a document somewhere the model's retrieval pipeline is configured to pull from, a knowledge base, an inbox, a shared drive. The model retrieves that document as context and treats its contents as instructions, executing them without any direct interaction from the attacker. The attack surface here is the data pipeline, not the user interface, and that distinction matters enormously for how defenses have to be designed.

EchoLeak proved this class of attack is not theoretical. Disclosed by Aim Security researchers in June 2025 and tracked as CVE-2025-32711 with a CVSS score of 9.3, EchoLeak was a zero-click indirect prompt injection against Microsoft 365 Copilot. A single crafted email, retrieved through Copilot's RAG context, caused the model to execute attacker-controlled instructions and exfiltrate chat logs, OneDrive files, SharePoint content, and Teams messages to a server the attacker controlled, without any user needing to click, open, or approve anything. The exploit slipped past Microsoft's XPIA classifier, its link redaction, and its Content Security Policy by routing through an allowlisted Teams image proxy. Microsoft deployed a server-side patch in June 2025 as part of that month's Patch Tuesday update, and no confirmed in-the-wild exploitation has been reported. What makes EchoLeak matter beyond its own patch cycle is that it stands as the first documented case of prompt injection weaponized for concrete data exfiltration in a production AI system, proving the RAG injection surface is a real, enterprise-scale risk and not a research curiosity. Lakera AI's Q4 2025 Agent Security Trends Report, covered by eSecurity Planet, found that system prompt extraction was the single most commonly observed attacker objective in enterprise environments during that quarter. This is active exploitation happening now, in production systems rather than academic papers.

Why the standard defensive advice (keep secrets out of prompts) is necessary but not the whole answer

OWASP's core recommendation, keep credentials and connection strings out of system prompts and never lean on the prompt as a security control, is architecturally sound, and nothing in the record above argues otherwise. If a secret is never placed in the prompt, it cannot be extracted from the prompt. That much is simple.

But what if the prompt has no secrets in it and still causes damage when it leaks? Practitioners raise a fair objection here. Effective customization of an LLM requires detailed behavioral instructions: persona constraints, tool configurations, workflow logic. Strip a prompt down to bare, generic language and the model's usefulness degrades, its outputs get less predictable, and none of that actually removes the disclosure surface, because the model can still be coaxed into revealing whatever is left. The advice to keep prompts secret-free is correct and necessary. It simply isn't sufficient on its own: the leakage mechanism described earlier is structural, a function of how the model processes tokens rather than what specific content happens to be sitting in the prompt.

Software-layer defenses built to patch this gap run into their own limits. Yang et al. found that existing defenses either fail to prevent leakage without degrading usability, or require constant updates to address new attack variants, so there is no stable resting point in the prompt-engineering layer by itself. The most promising technical answer to emerge from the research community is ProxyPrompt, presented in the ACL 2026 Findings by Zhuang et al., which replaces the system prompt with a functionally equivalent proxy carrying an unrelated semantic meaning. Content extracted from a system protected this way neither preserves the original meaning nor works as a valid instruction set anywhere else, and its reported protection rate substantially beats the next-best baseline. ProxyPrompt requires access to model weights, the ability to inject custom embeddings, and gradient computation for optimization, and those prerequisites are unavailable in the black-box API setting that covers most enterprise deployments. The strongest defense on paper is unavailable to the teams who need it most. Detection and monitoring have to carry weight that prompt design alone cannot.

A Layered Mitigation Approach for Practitioners

No single control closes a structural leakage surface, so the realistic answer is a stack: prompt design, runtime controls, and monitoring, each catching what the others miss.

Credentials, API keys, and connection strings belong in a secrets manager or environment variables accessed by the application layer, never typed into the prompt itself, so the model never sees them in plaintext at all. System prompts should stay as behaviorally minimal as the use case allows, because every additional instruction is additional reconstructable intellectual property sitting in the context window. Where detailed behavior really is unavoidable, it helps to separate concerns: a lean prompt for persona and hard constraints, with domain-specific knowledge delivered through structured retrieval instead of being written directly into the instructions. Salesforce AI Research tested a combined defense against the multi-turn sycophancy attack described earlier, pairing a query-rewriting defense at the first turn with an instruction-based defense at the second, and found it reduced the average attack success rate to a residual level, the best outcome achievable through prompt-engineering defenses alone in a black-box setting.

Runtime controls sit above the prompt layer and try to catch extraction attempts as they happen. Input validation can flag known extraction patterns, direct instruction-repeat requests, role-play prompts angling for elevated privilege, delimiter confusion tricks, before they ever reach the model. Output filtering does the mirror-image job, scanning responses for prompt fragments, credential patterns, or structural markers of system prompt content before anything gets returned to the user. In multi-agent systems specifically, permissions have to travel with every handoff between agents. Fiddler's 2026 analysis identifies privilege escalation through agent handoff as its own distinct leakage vector, one that demands explicit access-control propagation rather than an assumption that permissions carry over automatically.

Monitoring catches what slips past the first two layers. Fiddler's analysis names three real-time signals: PII detection rates in model output, shifts in output entropy that flag statistically anomalous response distributions, and anomalies in response length, all detectable without needing access to the model's internals. The gap between when a leakage event happens and when someone notices is where enterprise risk actually accumulates, and closing it takes continuous logging with near-real-time alerting rather than a quarterly audit. In agentic deployments especially, this kind of monitoring stops being a nice-to-have. When an agent retrieves external data, processes it, and acts on it autonomously, a single injected instruction can move through an entire workflow before any person ever looks at it.

How Agentic Architectures and Shadow AI Amplify the Leakage Surface

Everything described so far assumes something like a single model behind a single interface. Agentic deployments and unofficial, unmonitored AI tools break that assumption, turning a leak from a single point of failure into a systemic risk across the enterprise, because every new agent or shadow tool adds an injection surface that the existing controls were never built to watch.

In agent-based systems, prompt leakage can expose more than the prompt text itself. It can reveal backend API calls, implementation details, and the underlying system architecture, giving an adversary a much fuller picture of how the system actually operates. Help Net Security's March 2026 report on enterprise AI agent security found that a meaningful fraction of deployed agents can create and task other agents on their own. Each new agent spun up this way compounds the attack surface, and it does so without necessarily passing through whatever review process governed the original deployment.

The security community has started formalizing this shift. The OWASP Top 10 for Agentic Applications, released in December 2025, names Agent Goal Hijack as ASI01, its top-ranked threat in the agentic context, with prompt injection identified as the primary way that hijack gets delivered. That ranking is a recognition that the threat model for agentic systems has moved past what a single hardened prompt or a single output filter can cover. Static defenses built around one model, one prompt, and one interface were never designed for a world where agents spawn other agents, retrieve from pipelines nobody centrally tracks, and hand off permissions across boundaries that shift with every new integration. The architecture that makes agentic AI powerful also makes its leakage surface nearly impossible to fully map in advance, a reckoning every team building on this technology eventually has to have.

Sources

  1. How do you detect when an LLM agent is leaking system prompt content in its responses?
  2. Prompt Leakage effect and defense strategies for multi-turn LLM interactions
  3. Understanding and Mitigating Prompt Leaking Attacks in
  4. Understanding and Protecting Against LLM07: System Prompt Leakage
  5. EchoLeak: The First Real-World Zero-Click Prompt Injection Exploit in a Production LLM System
  6. ProxyPrompt: Securing System Prompts against Prompt Extraction Attacks
  7. Understanding and Mitigating Prompt Leaking Attacks in Real-World LLM-Based Applications
  8. System Prompt Extraction Attacks and Defenses in Large Language Models

More in Data Exfiltration