Security/ ai-agents · prompt-injection · computer-use-agents · security-research

A New Attack Tricks AI Agents Into Picking the Bad Option

A new benchmark shows AI agents that click and type for you can be steered into harmful actions without a single injected command.

A new study shows that AI agents which click, scroll, and type on your behalf can be manipulated into taking a harmful action - even when every word on the page looks harmless.

Researchers tested Computer Use Agents, AI systems that control a web browser or desktop to complete tasks, against a new benchmark called STEER-Bench: 101 tasks spanning nine domains. The attack, dubbed branch steering, doesn't inject rogue commands into a page. Instead, it nudges the agent toward a pre-approved action path - like clicking a legitimate-looking but harmful button - that the agent's own plan already allows. Against standard agents, the attack worked 94.4% of the time. Against Dual-LLM setups, an architecture that splits planning from reading untrusted content, it still worked 89.5% of the time. The researchers' own fix, called COBRA, pairs the trusted plan with hard limits on what parameters and destinations each branch can touch, cutting the attack's success rate to zero while keeping 97% of normal task performance.

That Dual-LLM number is the real story here. It's the closest thing the field has to a formal security guarantee for agents handling untrusted web content, and it was already treated as the strong answer to prompt injection. This paper shows that guarantee assumes a static plan, and GUI agents can't work that way - they have to branch based on what a page actually shows them, and that branching is exactly what attackers can steer.

Call it the GUI tax on agent security: every dynamic decision point is a door, and right now most of those doors don't lock. COBRA looks promising on paper, but it's a benchmark result, not a shipped product - expect this cat-and-mouse to run for a while before letting an agent click buttons unsupervised feels safe.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →