Est.

Least Privilege Design for Enterprise AI Agents

Stopping AI agent attacks means limiting what they can access in the first place.

Contributing Editor · · 10 min read
Cover illustration for “Least Privilege Design for Enterprise AI Agents”
AI Agent Security · October 3, 2026 · 10 min read · 2,233 words

Least privilege is not a new idea, and nobody seriously disputes it. Give a principal only the access it needs, for only as long as it needs it. The trouble with AI agents is that they are a different kind of principal than the one identity and access management was built to handle, and the old playbook doesn't map cleanly onto something that doesn't behave like a person.

Traditional IAM assumes a human sitting at a keyboard: one identity, one session at a time, a person who has to click "approve" before anything consequential happens. An AI agent can violate all three of those assumptions in a single workflow. CISA, the NSA, and the NCSC describe agentic AI through three properties that explain why: autonomy, statefulness, and tool access. An agent takes an underspecified goal and runs with it, keeps acting without continuous human intervention, and can spawn additional agents to finish sub-tasks on its own.

Large language models process instructions and data through the same channel, with no hard wall between the two. That means every document an agent retrieves, every email it reads, every web page it fetches, every message it gets from another agent, is a candidate instruction as far as the model is concerned. The trust boundary isn't fixed at design time. It expands, silently, to include every byte of untrusted content the agent happens to touch during its work.

That would be a containable problem if agents only returned information, the way a vulnerable web form might leak a database value to whoever asks. Agents don't just answer questions. Statefulness combines with this to increase the danger. A poisoned memory entry planted today doesn't need to be triggered today. It can sit quietly and execute against some unrelated query weeks later, long after anyone would think to look for it.

How agent over-privilege becomes a reachable attack surface

The logic leads to an obvious stake: if an agent can call an API it doesn't need, anyone who can influence that agent gains the same reach. The damage from a prompt injection or a credential leak doesn't scale with how clever the attacker is. It scales with how much access the agent was carrying from the start.

A survey from Obsidian Security found that 90% of agents hold excessive privileges. That means the access an agent's configuration describes and the authority it actually carries inside connected systems are, for nine out of ten agents studied, two different things. Excessive autonomy means high-impact actions happen without a human anywhere in the approval chain.

Because no defense against prompt injection works every time, the only approach that holds up is containment: assume an injection will eventually succeed, and design the system so that a compromised agent still can't do serious damage or reach sensitive endpoints. Prevention alone doesn't get you there.

Reusing a human user's session or sharing a single key across agents makes all of this worse at once. Microsoft's July 2026 security guidance names the organizational pattern that produces this directly: a team provisions a broad read-only role because the first use case looks read-only, the workflow later expands to include remediation, write access gets bolted on without anyone rethinking the role, and the expansion never gets revisited.

Even when each individual connection looks modest, the combination can create risk nobody signed off on. An agent wired into email, a file store, a ticketing system, and a code repository might look low-risk at each individual connection point. Strung together, that same agent can correlate data across all four systems and take actions that no single integration review ever considered as a whole.

What production exploits reveal about permissive agent scope

None of this is hypothetical anymore. A handful of documented incidents show how over-privileged agent scope turns into a working attack path, and the scale of the damage tracks the breadth of what the agent could reach, not the sophistication of whoever found the opening.

EchoLeak, tracked as CVE-2025-32711 and disclosed by Aim Security researchers in June 2025, carries a CVSS score of 9.3, critical. It stands as the first documented case of prompt injection weaponized for actual data exfiltration against a production AI system.

What made EchoLeak dangerous wasn't the injection technique alone. It was how much Copilot could reach once the injection worked. A narrowly scoped agent hit with the same attack would have handed over far less, simply because there would have been far less within reach. That's the least-privilege argument made in the field rather than on paper. Microsoft patched the specific vulnerability server-side, but indirect prompt injection is a structural property of how retrieval-augmented AI assistants work, and no single fix retires the category.

A second case, known as BodySnatcher, hit the ServiceNow AI Platform. Once inside, they could invoke AI workflows and plant backdoor accounts carrying elevated privileges. The identity model being too loose is what turned a login bypass into full workflow execution.

A third incident touched OpenAI in May 2026, through a supply chain compromise. The compromise happened where the agents involved kept functioning exactly as intended, from the operators' point of view. The break was at the credential layer, not inside the model itself.

The three incidents line up into a pattern. Each exploit worked by inheriting or stealing broad access that was already sitting there, waiting, rather than by defeating some clever model-level defense. That's what makes identity design the place to intervene, and it's where the rest of this piece turns.

Giving each agent a distinct, lifecycle-managed identity

Every one of those incidents traces back to an identity that was shared, borrowed, or loosely defined. Fixing that starts with treating each agent as its own principal. Each agent needs an identity of its own: a distinct identity with a stated purpose and a human owner behind it, rather than a shared secret or a service account recycled across five different workflows.

Microsoft's July 2026 guidance frames the governing idea this way: treat every agent as a first-class principal. Give it an identity that gets managed through its full lifecycle, from creation to retirement. Assign it explicit roles. Scope its permissions tightly. Limit what tools it can touch to a preconfigured manifest rather than whatever happens to be reachable on the network.

What does a dedicated per-agent identity actually look like in practice?

The cost appears later, usually during an investigation nobody wanted to run. Logs might capture which tool got called and when. They can't answer who authorized the action, under what role, or whether the action fell inside the agent's intended scope, because the identity model itself never defined roles, authorization, or scope clearly enough to make that question answerable. Whether an agent is acting under its own identity, under a human's delegated session, or some blend of the two has to get settled when the system is designed. Discovering the answer mid-incident is too late.

The same discipline has to extend down through multi-agent architectures. When an orchestrator spawns a sub-agent to handle one piece of a task, that sub-agent needs its own scoped identity rather than simply inheriting whatever credentials the orchestrator was carrying. If that sub-agent simply inherits the orchestrator's credentials instead of getting its own scoped identity, the over-privileging persists, reproducing itself one layer down, at every branch of the chain.

Scoping roles and credentials to the minimum necessary for each task

Identity answers who the agent is. Scope answers what it's allowed to do, and for how long. Both matter, but scope is where the actual permission boundaries get drawn, and it's where most of the incidents above found their opening.

Build roles around the smallest meaningful unit of work rather than around a department or a project. What should be avoided is bundling a handful of unrelated permissions into one composite role just because the agent might eventually need all of them.

Workflows that combine evidence-gathering with remediation deserve particular care. The instinct to fold read access and write access into a single role, because the workflow is going to need both eventually anyway, is the scope-creep pattern that shows up in Microsoft's own case study. The fix is to separate those duties and put a step-up approval in front of anything with a write effect.

One detail changes who gets to make the authorization call. A model can be talked into almost anything given the right prompt. A policy engine sitting outside the model, evaluating requests against fixed rules, cannot be talked into anything.

One might reasonably push back here. Doesn't scoping credentials this tightly, and forcing step-up approvals at every write boundary, slow the agent down? It does add friction, and that's a fair trade-off to weigh rather than wave away. But treat that friction as an engineering constraint to design around, the same way latency or rate limits are, rather than a reason to hand an agent standing privilege it only needs occasionally. The deeper tension between speed and oversight gets its own treatment further down.

Binding tools to an explicit allowlist and enforcing it at the policy layer

Scoping an agent's role is necessary work, but it isn't sufficient on its own. A tightly scoped role doesn't help if the agent can still reach any tool sitting on the network. The tool ecosystem needs its own boundary, enforced outside the model, separate from the role-based access control layer.

Every agent-accessible tool must sit behind an API gateway or a policy enforcement point, which validates each request against authentication, authorization, schema enforcement, parameter constraints, quotas, and rate limits. The allowlist itself is the control. The absence of an explicit deny rule is not a substitute for one.

Microsoft's July 2026 guidance frames this as a tools manifest, a version-controlled, declared artifact that spells out exactly which tools an agent can use, rather than leaving tool scope as an implicit side effect of whatever APIs happen to be network-reachable. A manifest can be reviewed, diffed, and rolled back. An implicit set of reachable endpoints can only be discovered, usually after something has already gone wrong.

Privilege escalation through tool chaining is the specific risk this guards against. Strung together, they're an exfiltration. That chain only works if the tool policy fails to check the second call independently of the first. Per-tool authorization, evaluated fresh at each step, breaks the chain regardless of what came before it.

Recent research gives a concrete picture of what principled enforcement at this layer looks like. The authors call the resulting property monotonic confinement: left to its own devices, the agent's effective action space can only shrink, never grow, which closes off silent privilege escalation even when an attacker is actively trying to manipulate the agent's behavior.

One layer gets missed often enough to call out directly: memory. Vector databases, conversation history, and session state are tools the agent reads from and writes to, exactly like any external API. The same allowlist discipline that governs outbound calls needs to cover what the agent is permitted to write into its own persistent memory. Memory poisoning, planting a false or malicious entry that influences a future decision, is its own attack path, separate from anything that happens through an external tool call, and it deserves its own line of defense rather than an assumption that the tool allowlist already covers it.

Where human approval gates belong in an autonomous workflow

Excessive autonomy is the third root cause in OWASP's breakdown of Excessive Agency, alongside excessive functionality and excessive permissions, and it's the hardest of the three to fix with a configuration change. Scoping roles and binding tools reduces what an agent can reach. Neither one answers the separate question of when a human needs to be in the loop before an action actually runs.

Sign-off is required only for certain actions, not for every one. A blanket rule requiring sign-off on every action would make an autonomous agent pointless. The gate belongs specifically in front of actions that are irreversible, disproportionate in their consequences, or hard to unwind once they've happened: a wire transfer, a production database write, a change to an IAM policy, the creation of a new account. Read-only retrieval and draft-stage work, the kind of role scoping described earlier, don't need the same friction, and treating them as though they do just trains people to click through approvals without reading them.

The line sits where it does because of the reversibility and blast radius of the specific action in front of the agent, not its general trustworthiness. A query against a production database and a deletion of records off that same database carry entirely different risk profiles, even though both might get routed through the same agent, using the same credentials, in the same session. Treating them identically, either by gating both or by gating neither, misses the distinction that actually matters.

That's the throughline across every control this piece has walked through. Distinct identity tells you who acted. Scoped roles tell you what they were allowed to do. Tool allowlists tell you where they were allowed to go. Approval gates tell you which of those actions needed a person watching before it happened. None of the four substitutes for the others, and skipping any one of them is how a well-intentioned agent ends up with the kind of reach that turned EchoLeak from a curious bug report into a critical, exploited vulnerability.

Sources

  1. Least privilege for AI agents: Identity, access, and tool binding
  2. AI Agent Security Checklist (2026): Agentic Risks & Controls
  3. Securing Agentic AI in the Enterprise: A 2026 Practitioner Guide
  4. Careful adoption of agentic AI services
  5. Parallax: Why AI Agents That Think Must Never Act

More in AI Agent Security