Researchers built a policy engine that keeps compromised AI agents from going rogue - and it works most of the time.
The system, called Sapien, is a policy engine that sits between an AI agent and the tools it can call. Instead of a static allowlist of approved actions, Sapien tracks what the agent has already done and uses that history to decide what it is allowed to do next. Policies are written as regular expressions extended with stateful predicates, checks generated on the fly, and scoped rules that only apply in certain contexts. The researchers tested it on two benchmarks, AgentDojo and Toolathlon, simulating agents that have been completely taken over by an attacker.
The results matter because most AI agent security today still assumes the agent itself can be trusted to behave, or relies on fixed lists of permitted tools that do not account for multi-step tasks. Sapien blocked 93 to 95 percent of attacks on AgentDojo and 62 to 85 percent on Toolathlon, roughly double what simple tool allowlists managed on longer tasks, while costing only a few percentage points of the agent's normal usefulness.
That is a real improvement, not a guarantee. Even the strongest result here still lets one in twenty attacks through, and the weaker Toolathlon numbers mean nearly four in ten attacks can still slip past on harder, longer tasks.