Security/ ai-agents · security · dev-tools

Study Finds AI Coding Agents Can Silently Swap Approved Actions

A new study shows popular AI coding agents like Claude Code and Cursor can execute different actions than the ones a human actually approved.

Researchers found that the approval screen in AI coding agents doesn't always mean what you think it means.

A new arXiv paper examines "approval laundering" in AI coding-agent harnesses like Claude Code, Codex CLI, and Cursor - cases where the action a human approves isn't the action the system actually runs. The authors identify six distinct failure modes (Scope, Argument, Temporal, Tool, Delegation, and Semantic laundering) and test all of them in a controlled, repeated study on Claude Code's pre-execution checkpoint. They also build a prototype fix, called an Approval Token, that cryptographically binds an approval to a specific agent, session, and set of arguments before anything runs. In testing across 118 paired runs, the token eliminated two of the six failure modes outright but left two others completely unaffected.

These harnesses are built to let an AI agent write and execute code with a human in the loop, and the approval prompt is the entire safety mechanism, not just a formality. The study's most notable finding is a negative one: scope and argument laundering survive the fix because they leave every recorded field of the dispatch unchanged, diverging one process layer below what the harness's own checks can observe.

A security model is only as strong as the gap between what it asks and what it does, and this study spells out exactly how wide that gap already is.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →