A new paper proposes a fix for AI agents that approve a tool call, then execute it only after the permissions behind that approval have quietly changed.
The method, called BSC-R, locks a one-time commit authorization to both the specific action and the exact state of the world that justified it, so a stale or altered authorization cannot be reused. Tested on 2,847 attacked AgentDojo episodes, it preserved the agent's normal task performance (80.576% utility) while cutting successful attacks to 1.616%. Across 10,302 frozen proposals, it approved every legitimate unchanged commit and blocked all ten categories of altered or replayed ones. A separate boundary-drift test pushed further: it rejected all 4,403 invalid contexts while still approving all 5,899 valid ones.
This targets a specific, underappreciated failure mode. Agents with real tool access, like email, file systems, or payments, often check permissions once and then act on stale information, which is the agentic equivalent of a classic time-of-check-to-time-of-use bug. As more products hand language models write access to real systems, that gap between deciding and doing is exactly where an attacker would look to slip in a change.
The honest part is what happens next. On a prospective, independently run evaluation, the 3,460-scenario CONTINUITY suite, BSC-R handled the easy cases cleanly: all 700 benign runs, 1,200 of 1,200 non-replay attacks, and all 160 replay and 200 ambiguous cases. But across that suite's full 2,560 attacks, its invalid-effect commit rate was 25%, compared with 0% for CONTINUITY. A method that looks airtight on its own tests still let a quarter of bad commits through once it met a suite it did not build.