Est.

Supply Chain Risk in AI Agent Tool Registries

Malicious skills poisoned AI agent registries before defenses could catch up.

Senior Writer · · 8 min read
Cover illustration for “Supply Chain Risk in AI Agent Tool Registries”
AI Agent Security · October 4, 2026 · 8 min read · 1,902 words

AI agent tool registries have built the same kind of ecosystem that npm and PyPI built for software packages, and they've inherited the same weakness along with it. A CSA research note draws that line directly: npm and PyPI let software grow fast, but that speed came with an attack surface that took the software world years to even begin to lock down. Agent skill marketplaces grew fast too, but you won't find anywhere near that amount of hardening behind them.

The infrastructure tying agents to tools mostly runs through one standard now: MCP, the Model Context Protocol, has become the default way agents connect to outside tools. ChatGPT's plugin ecosystem, autonomous coding frameworks, and enterprise workflow automation tools all share this same basic setup underneath, even when they look different on the surface. One agent calls out to many tools, and each tool gets to describe itself to the model in its own words.

That's exactly where things get fragmented. Researchers behind the AgentHub project, out of Loyola University Chicago, Purdue University, Socket Inc., and Oracle Labs, describe the infrastructure for discovering, evaluating, and governing agents as something that "remains fragmented compared to mature ecosystems like npm." Naming conventions exist. Distribution channels exist. Protocol descriptors exist. What doesn't exist yet, at any real scale, is a registry layer that makes agents "discoverable, comparable, and governable."

Why does that gap matter more here than it would in a normal software registry? The thing moving through an agent registry isn't a binary, but instructions that the model reads, trusts by default, and treats as guidance for what to do next.

Agents aren't passive text generators sitting behind a chat window. Stellar Cyber's late-2026 threat briefing frames the shift this way: agents execute code, modify databases, call APIs, and hold onto long-term memory without a person checking each step. That's a move from content generator to active participant in enterprise infrastructure. A compromised npm package has to get executed before it does damage. A compromised skill file just has to get read by a model that was built to follow what it reads. The code supply chain problem is old news. The natural language supply chain problem is just getting started, and it runs on trust the model was designed to extend.

The ClawHub campaign and a poisoned skill registry at scale

The ClawHavoc campaign turned that structural risk into a documented event. It showed that an attacker doesn't need to find a bug in the agent framework itself to compromise a skill registry at scale. The installation pathway was the vulnerability.

In late January and early February of 2026, a researcher using the handle depthfirst disclosed CVE-2026-25253: a one-click remote code execution flaw in OpenClaw, reached through stealing authentication tokens. It's among the first CVE identifiers ever assigned to an agentic AI system, the SkillFortify formal analysis paper found. Within days of that disclosure, the ClawHavoc campaign flooded ClawHub, OpenClaw's skill marketplace, with malicious skills. Koi Security researcher Oren Yomtov identified a substantial tranche of malicious skills among those listed at the time. A security vendor's analysis and another research team's review both found that a significant share of the marketplace's historical catalog was malicious.

How did the malicious skills actually work? The instructions lived inside SKILL.md configuration files, plain natural language files that tell an agent when and how to use a given skill. Attackers used those files to guide the agent into presenting fake prerequisite installation steps to the user. Following those fake steps led to execution of AMOS, a credential-stealing malware family built to pull saved passwords, browser data, and crypto wallet files off a compromised machine. None of it tripped conventional file-based detection, because none of it looked like a file doing something malicious. It looked like an agent walking a user through a normal setup step.

This wasn't an isolated flare-up. The SkillFortify paper notes that a tool called MalTool has synthesized 1,300 standalone malicious tools, on top of thousands of additional tools carrying embedded malicious behavior aimed at LLM-based agents. The same paper found that VirusTotal fails to catch most agent-targeted malware. ClawHub wasn't a one-time lapse in an otherwise clean system. It was the first big, visible instance of a detection gap that cuts across the entire category.

The SKILL.md file as the actual attack surface

Diagram: How a Poisoned SKILL.md Spreads Through Agent Context. Visualizes: Illustrate the four-stage infection chain documented in the ClawHub/ClawHavoc campaign: (1) a malicious SKILL.md is loaded into the agent's context window, (2) the agent…

The malicious packages in ClawHub mattered less than what they revealed about where the real risk sits. The attack surface is the configuration file that tells the agent what to do, because that file carries real operational weight once the agent reads it.

A SKILL.md file has no built-in limits on what it can say. It's a natural-language instruction set, and per the CSA CISO Briefing, once it's loaded into an active agent session, its content becomes part of the agent's working context. From there, it can shape tool use, file access, command execution, and network activity, the same categories of behavior a piece of executable code would touch, just without ever compiling or running as code.

Snyk's ToxicSkills audit, looking at the ClawHub corpus, sorted the malicious payloads it found into four families. Credential exfiltration instructions go after API keys and environment variables sitting on the machine. Command execution instructions push the agent to run shell commands, dressed up as a legitimate tool call. Data routing instructions quietly reroute the agent's output to a server the attacker controls, so the user never sees the result. Persistence instructions tell the agent to write more malicious instructions into other context files already on the local filesystem.

That last one deserves the most attention. Once loaded, a single poisoned skill can plant its own instructions into other files the agent will read later, so it spreads through the context the agent trusts rather than through any file-copying mechanism a scanner might catch.

The CVE record attached to this problem shows Embrace the Red publicly disclosed, on February 11, 2026, a technique for hiding adversarial instructions inside visually clean skill files, using invisible Unicode Tag characters in the range U+E0000 to U+E007F; a human reading the file sees nothing, while a model parsing it reads a full set of hidden commands. Anthropic patched Claude Code against it a day earlier, on February 10, 2026, and now detection refuses context files that contain those characters. The fix closed one specific trick. It didn't close the underlying gap: nothing in most agent platforms can tell a trusted developer's instructions apart from a third-party skill's instructions once both are sitting in the same context window.

MCP's extension of the same attack surface to any connected tool

ClawHub was one marketplace with one skill format. MCP is a protocol, adopted far beyond any single ecosystem, and it carries the exact same weakness built into its architecture. MCP loads third-party tool descriptions straight into the model's context window, word for word, just like a SKILL.md file does. Any tool an agent connects to through MCP inherits that exposure, whether it came from a curated marketplace or a developer's personal GitHub repo.

CVE-2025-54136, known as MCPoison, shows how that plays out in a real product. The flaw sits in Cursor IDE: once a user approves an MCP server configuration, Cursor trusts future entries under that same name without re-checking them. An attacker who controls a later update can silently swap out the server's commands or tool definitions. No re-approval prompt. No warning to the user. The door that was checked once stays open for anything walked through it later.

The deception runs deeper than the swap itself. A malicious MCP server can hide instructions inside a tool's description that never appear in the IDE's interface, because the UI renders only a simplified summary there instead of the full text. The model, though, reads the full description every time. A poisoned tool doesn't even need to be called to do damage. Its description alone can instruct the model to go pull SSH keys, configuration files, or anything else reachable through the other tools that agent happens to have access to.

None of these weaknesses sit in isolation. API security gaps, prompt injection, supply chain compromise, and session management flaws interact and build on each other across an MCP deployment. If a supply chain attack gets a malicious MCP server installed, it opens the door to tool poisoning. A session hijacking attack enables token reuse. A tool poisoning attack exfiltrates credentials that then enable the next stage of compromise. Each weakness hands the next one a foothold.

Making this worse is how the MCP registry ecosystem still largely operates on name recognition. Vetting is minimal. Developers often install a server because the name looks familiar, not because anyone checked the publisher's identity or read the source code behind it.

What does this mean for the person defending the network? The blast radius is every system the compromised agent can reach, not the tool registry itself. Stellar Cyber describes this as a confused deputy problem: the attacker doesn't need to breach the network directly. The attacker only needs to trick an agent that's already trusted inside that network into acting as the one who carries the data out.

Why scanners cannot see the most dangerous registry attacks

Antivirus scanning, code signing, dependency analysis, vulnerability databases: these are the tools an enterprise security team already has in place, and CSA's June 2026 research found every one of them was built to inspect executable code. None of them were built to evaluate natural language instructions. The sharpest attacks in this category don't trip any of them.

ClawHub didn't ignore the problem after ClawHavoc. The marketplace added VirusTotal integration and a proprietary scanning capability called ClawScan. Unit 42 researchers tracked the marketplace from February through May of 2026, and malicious skills still got past both scanners. One technique simply inflated file size past the point where the scanners would even finish processing it. Other skills carried no detectable payload at all, because there was nothing shaped like malware for a scanner to find.

A May 2026 academic study, posted as arXiv:2605.11418, put numbers to the gap. Metadata alone, with no code payload whatsoever, can manipulate which skill an agent chooses at a high win rate in head-to-head comparisons, bias agent selection in a majority of trials, and slip past automated governance classifiers across a wide range of conditions. None of it needs a single line of executable code.

Semantic Compliance Hijacking, documented in that same CSA research, makes the mechanism explicit. SCH writes malicious behavior directly into natural language, framed as a compliance requirement, an operational instruction, or a safety rule sitting inside a skill's configuration metadata. Static analysis has nothing to flag, because there's no executable code present to analyze. Pattern detection has nothing to match, because no code signature resembles known malware. CSA's research reports that SCH achieves complete evasion of scanner detection, a 0% catch rate against the tools currently deployed to catch it [5][6][7][8].

The SkillFortify formal analysis backs this up from a different angle. Every reactive security tool tested against the problem, including a widely starred open-source project that uses YARA pattern matching to scan skills, shares the same blind spot: they're built to recognize patterns in code, and the attacks that matter most here were never written in code to begin with. That's the gap a credible vendor AI risk program has to start from, not patch over.

Diagram: Why Scanners Miss the Most Dangerous Agent Attacks. Visualizes: Show the detection gap using the three concrete findings from the article: (1) ClawHub's VirusTotal integration plus proprietary ClawScan both failed — malicious skills evaded…

Sources

  1. Poisoned Skills: AI Agent Marketplace Supply Chain Attacks
  2. Formal Analysis and Supply Chain Security for Agentic AI Skills
  3. CISO Briefing: Agent Context Poisoning — SKILL.md and the New AI Supply Chain Attack Surface
  4. AgentHub: A Registry for Discoverable, Verifiable, and Reproducible AI Agents
  5. Agent Context Poisoning: SKILL.md and the New AI Supply Chain Attack Surface
  6. Poisoned Pipelines: Malicious AI Model and Skill Repositories
  7. Under the Hood of SKILL.md: Semantic Supply-chain Attacks on AI Agent Skill Registry
  8. AI Coding Agents: An Unaudited Supply Chain Node

More in AI Agent Security