Est.

Multimodal Prompt Injection via Image and Audio Inputs

Attackers can hide commands in images and audio that AI models execute without catching them.

Correspondent · · 10 min read
Cover illustration for “Multimodal Prompt Injection via Image and Audio Inputs”
Prompt Injection · September 30, 2026 · 10 min read · 2,219 words

A vision-language model doesn't read an image the way a person does. It folds the whole picture into a stream of tokens, the same kind of tokens it uses for text, and hands that stream to the same reasoning process that follows your typed instructions. That single design choice is why a picture can carry a command, why a snippet of audio can carry one too, and why the fix isn't a patch.

How multimodal models treat image and audio inputs as executable instructions

Vision-language models don't hold a mental line between "content the user is showing me" and "instructions I should follow." Cloud Security Alliance's research note notes that every pixel gets folded into contextual information the model uses to generate its response. That single design choice is the design itself, not a rough edge in an otherwise careful one.

Newer architectures make the problem sharper, not softer. NVIDIA's AI Red Team found that early fusion models, Meta's Llama 4 among them, blend text and vision tokens together from the start, so an emoji sequence or a rebus puzzle can act as a command with zero explicit text prompt attached. Defenses built to scan for suspicious text never even get a chance to look, because there's no text to scanc3.

Audio runs the identical playbook through a different sense. Speech recognition transcribes whatever it hears, spoken commands embedded in a clip included, and routes the transcript into the same pipeline as a legitimate spoken request. Nothing in that pipeline is built to ask whether the voice giving the command was supposed to be there.

That leaves a structural hole. Input sanitization, keyword filters, XSS filters: all of it works on characters, on text strings, on the stuff you can grep. None of it looks at a pixel grid or a waveform for hidden meaning, so image and audio inputs pass straight through unguarded. Researchers have started calling the resulting mess "Prompt Injection 2.0," a label for the moment natural-language injection tactics fused with multimodal exploits into an attack surface nobody built the earlier defenses to catch.

The four image injection classes

Once you accept that images can carry instructions, the next question is how attackers actually encode them. Cloud Security Alliance's note maps four separate techniques, and a defense tuned for one often does nothing against another.

Typographic injection hides instructions as visible text: low contrast, tiny fonts, rotated lettering, text dropped into a busy background where a human eye skips past it but the model reads it fine. The CSA note recorded substantial success rates in black-box tests against GPT-4V, Claude 3, Gemini, and LLaVA, even under conditions built to stay stealthy.

Steganographic injection skips visible text. Instructions get encoded through least-significant-bit tricks, frequency-domain transforms, or a learned neural encoder, and the resulting image looks identical to the original to any person checking it by eye. Human review, as a checkpoint, does nothing here.

Adversarial perturbation works at the pixel level: noise a person can't perceive but that's been mathematically shaped to steer the model's output. The CrossInject framework, shown at ACM MM 2025, paired Visual Latent Alignment with Textual Guidance Enhancement run through surrogate open-source LLMs, and it pushed attack success rates well past earlier perturbation methods. Because the noise pattern can be tuned against a specific detector, this class stays dangerous even as defenses improve, since the attacker just retunes the noise.

Physical-world signage takes the same logic off the screen. Misleading text or objects placed in the real environment get read by embodied AI systems and acted on directly. CHAI researchers from UC Santa Cruz and Johns Hopkins, presenting at the 2026 IEEE Conference on Secure and Trustworthy Machine Learning, hijacked drones, autonomous vehicles, an aerial tracking agent, and an actual robotic vehicle using nothing more exotic than adversarial signage in the physical world.

None of these four sit in isolation. Chain-of-Attack research from Xie and colleagues, presented at CVPR 2025, found that stacking techniques, steganographic embedding layered with semantic manipulation for instance, beats either technique running alone. The taxonomy is a map of how attacks escalate when someone bothers to combine them.

Audio inputs extend the same attack logic into acoustic channels

When the pixel grid is replaced with a waveform, the same architectural flaw appears in audio, only harder to catch. Attackers can bury commands inside a podcast, a video's soundtrack, or background audio using psychoacoustic masking or ultrasonic carriers pitched above what a person notices, and a speech-recognition system transcribes and acts on that hidden command without flagging anything unusual.

Spoken injection carries a stealth advantage image attacks don't get. A visual overlay can look off, a font rotated at a strange angle, a patch of pixels that doesn't quite match the lighting. Spoken content, on the other hand, is a native part of any video's audio track, and the transcription systems models get trained on treat narration as trustworthy by default. Misleading speech doesn't stand out the way a doctored image sometimes does.

Enterprise voice systems carry this risk directly into daily operations. Any call-center IVR system handling customer records is exposed the moment an attacker manages to slip adversarial audio into hold music or ambient background sound.

The audio injection has not been studied as thoroughly as image injection. The research base on end-to-end audio attack mechanisms is thinner, and attack-defense parity for audio hasn't been established the way it has for images. The risk is documented and real, but its full scope is still being mapped, not fully known.

Coordinated cross-modal attacks multiply success rates beyond any single channel

Attacking one channel at a time is the easy case. Attacking two at once, and lining them up so they agree with each other, is where the numbers turn ugly.

Audio and video get processed through separate perceptual pathways inside a multimodal model, but both feed into the same downstream reasoning step that produces the final output. Audio and visual modalities are processed through distinct perceptual pathways in MLLMs, but both feed the same downstream reasoning and output generation, so when adversarial signals arrive through both channels simultaneously and are semantically aligned, the model has no independent channel to use as a cross-reference. There's no built-in skepticism, no cross-reference, no second opinion. Whatever the aligned channels agree on, the model tends to follow.

Content moderation turns out to be a particularly soft target for this. Injecting safe-sounding speech into a video that's visually harmful measurably weakened a model's ability to flag the harmful content. Cross-modal injection can knock out the safety classifier meant to catch the bad content.

A real-world test made the stakes concrete rather than theoretical. That overwrite opened the door to remote code execution and secret exfiltration down the line. One image, one file overwrite, and the agent's future behavior was compromised before anyone noticed.

Agentic pipelines and production deployments face the worst exposure

Attack success rates in a lab are one thing. What happens when that success rate meets an autonomous system running unattended in production is another matter entirely, and it's a worse one.

Agentic AI systems that pull in images, audio, and documents from sources nobody vetted are the most exposed setup around, because one injected input can ride through an entire multi-agent workflow and touch everything downstream of it. OWASP already ranks prompt injection, listed as LLM01, as the top-severity risk in production LLM deployments, and the 2025 revision folded multimodal cases, instructions hidden in images specifically, into that same top-tier classification. Regulators and compliance teams are treating this as a first-order production risk, not an edge case.

The exposure doesn't stop at the chatbot window. Agentic AI, RAG pipelines, multimodal models, and AI coding assistants each open their own injection vectors, and none of them get covered by defenses built for plain text.

The evidence isn't hypothetical. In August 2025, researcher Johann Rehberger filed prompt injection CVEs within a single month against six production tools, not six research demos: GitHub Copilot, Claude Code, Cursor IDE, AWS Kiro, Google Jules, and Amazon Q Developer. Several of the resulting CVEs scored severely on the CVSS scale: Microsoft Copilot at 9.3, GitHub Copilot at 9.6, Cursor IDE at 9.8. Those numbers sit near the top of the scale, reserved for flaws that let an attacker do serious damage with little effort.

Earlier in 2025, Rehberger showed something arguably stranger with Google's Gemini Advanced. Uploading a document containing hidden prompts caused Gemini to store fabricated personal details in its long-term memory through delayed tool invocation, and those fabricated details then shaped every subsequent session. The injection didn't just corrupt one conversation. It planted a false memory the model kept carrying forward.

Regulators have taken notice on their own timeline. Prompt injection maps to at least seven major frameworks, MITRE ATLAS, NIST, EU AI Act, ISO 42001, GDPR, and NIS2 among them, and the EU AI Act's August 2026 Article 50 transparency deadline, with high-risk system obligations now deferred to December 2027, makes compliance mapping urgent for organizations deploying multimodal systems. Any organization running multimodal systems has reason to map its compliance exposure now rather than later.

What current defenses can and cannot do across image and audio vectors

Defenses exist. None of them close the door all the way, and the strongest evidence for that comes from researchers who tried to break their own field's best work.

"The Attacker Moves Second," a paper from Nasr and colleagues published in October 2025 with researchers from OpenAI, Anthropic, and Google DeepMind, tested twelve published defenses against adaptive attackers using gradient descent, reinforcement learning, random search, and human red teaming. Every one of the twelve defenses fell, most at high success rates for the attacker. That means treating any single published defense as temporary, since an adaptive attacker who studies the defense generally finds a way around it.

Commercial tooling hasn't caught up to the threat it's supposed to cover. As of April 2026, most commercial detection products, Lakera Guard, LLM Guard, and Azure Prompt Shields among them, still operate on text only in their public APIs. That leaves image and audio injection completely outside the coverage of the tools most organizations have already bought and deployed.

Two research efforts show what a more serious attempt looks like, and both come with an honest limitation attached. CaMeL, from Google, Google DeepMind, and ETH Zurich, wraps a protective system layer around the LLM so the system stays secure even if the underlying model itself can still be fooled. It holds onto meaningful task completion compared to a totally undefended baseline, trading a modest amount of usefulness for a provable security guarantee. That trade-off, quantified rather than hand-waved, is the clearest example available of what security actually costs in practice.

SUAD, a 2025 defense built specifically for audio, trains against ultrasonic attacks by randomizing both the timing and frequency of the attack signal, then generates a suppression signal that blocks the attack while letting real voice commands through. SUAD reports defense success above 98% against three separate types of ultrasonic attack. Strong numbers, but the scope stops at the acoustic channel. It says nothing about images, video, or the physical-world cases described above.

LLM-as-judge detection, where a second model checks the first model's inputs for semantic tricks a pattern filter would miss, is currently the only real option for catching subtle semantic injection. It also comes with latency and cost that make it a hard sell for anything running at scale. The CSA research note recommends defense-in-depth: input-layer detection, architectural privilege minimization, runtime behavioral monitoring, and mandatory human-in-the-loop gates for high-stakes actions. No one layer does the job alone.

Diagram: Twelve Defenses, Zero Survivors. Visualizes: Show the outcome of the 'Attacker Moves Second' study (Nasr et al., October 2025, with researchers from OpenAI, Anthropic, and Google DeepMind): twelve published defenses were tested against…

What modality-specific and continuous monitoring must address beyond generic governance

A checklist filled out once, at the moment a vendor gets approved, tells an organization nothing about how that same model behaves six months later against an image or audio input nobody tested. Multimodal injection exploits channels text-layer tools were never built to inspect, and the only workable response pairs modality-aware assessment of what's actually coming into the model with continuous monitoring of how AI systems behave once they're live.

The target keeps moving, too. Every time a model adds a new sense, vision, audio, video, physical-world sensing, it opens a fresh injection channel that didn't exist when the vendor was first evaluated. A point-in-time audit goes stale the instant a vendor pushes an update to how their model handles input, and there's no way around that except watching continuously rather than checking once.

The physical world makes the argument impossible to wave off as academic. The PI3D paper from 2026 and the CHAI work on autonomous vehicles and drones both show the threat has moved off the screen and onto the street. An organization running embodied AI or a real-time audio agent can't lean on network-perimeter security to catch an attack that arrives through a camera or a microphone, since perimeter controls were built to watch network traffic, not pixels or sound waves.

Video-capable models inherit the worst of both worlds. Video-capable VLMs inherit the vulnerabilities of both visual and audio channels plus unique risks that emerge only when the channels are combined, and any organization evaluating a video-capable system needs to test it as its own category, treating it differently from an image model with a soundtrack bolted on.

Sources

  1. Image-Based Prompt Injection: Hijacking Multimodal LLMs Through Visually Embedded Adversarial Instructions – Lab Space
  2. A Systematic Study of Cross-Modal Typographic Attacks on Audio-Visual Reasoning
  3. Multimodal Prompt Injection Attacks: Risks and Defenses for Modern LLMs
  4. Piggybacking on Perception: Stealthy Concurrent Audio Prompt Injections against Multimodal LLM Agents
  5. Image-based Prompt Injection: Hijacking Multimodal LLMs through Visually Embedded Adversarial Instructions
  6. LLM01:2025 Prompt Injection - OWASP Gen AI Security Project
  7. [2504.14348] Manipulating Multimodal Agents via Cross-Modal Prompt Injection
Filed underPrompt Injection

More in Prompt Injection