Security/ prompt-injection · ai-security · llm-agents · activation-steering

New Defense Blocks Most Prompt Injection Attacks on AI Agents

A new technique steers an AI model's internal activations during inference, suppressing indirect prompt injection without fine-tuning or detection logic.

A research team has found a way to make AI agents largely ignore malicious instructions hidden inside the documents and web pages they read.

The technique, called CounterSteer, works by finding a direction inside a model's internal activations - its "residual stream" - that corresponds to following an embedded instruction, then subtracting that direction from every piece of retrieved text before the model processes it. No fine-tuning, extra model, or added tokens are required, just access to the model's internals and knowledge of where tool-result text starts and ends. Tested across five open-weights models ranging from 8 billion to 106 billion parameters and five different vendor lineages, the fix cut attack success rates from a range of 21-100% down to 0-17%. On the AgentDojo benchmark, the rate of attacks that actually hijacked an agent's task fell from 10-49% to under 8%, while legitimate task performance held at 93-100% of baseline.

Indirect prompt injection - instructions smuggled into search results, emails, or scraped pages - is one of the more practical ways to hijack an AI agent, because the attacker never touches the prompt directly. Most defenses either filter inputs, which attackers can dodge, or retrain the model, which is expensive and still imperfect; this one edits behavior at inference time, with no separate detection step for an attacker to evade. The researchers are upfront that it is not a complete fix: when an attack manipulates the arguments of a legitimate tool call rather than hijacking the task itself, the defense only resists about 13 of 18 tested cases.

Translation: your AI agent is harder to trick into doing something it wasn't asked to do, but if an attacker just nudges the numbers in a call it was always going to make anyway, it is still largely on its own.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →