Researchers built a benchmark that asks not just whether an AI agent's safety monitor catches misuse, but whether it catches it in time.
The team released a set of roughly 6,200 transcripts testing agent monitors against two separate threats: prompt injection, where a compromised tool slips in a malicious instruction, and decomposition attacks, where a harmful request gets split into innocent-looking sub-steps. Each transcript is labeled with a harm window, the span between when the agent first commits to a harmful action and when it finishes carrying it out. Across 17 monitor setups, monitors that watch the agent's actions rather than its raw text scored well on both threats, with an AUC of 0.95 on decomposition and 0.99 on injection. Monitors that only read conversation content collapsed on injection attacks, scoring just 0.52, barely better than a coin flip.
The more interesting finding is about timing, not accuracy. Even the strongest monitors struggled to flag decomposition attacks while they were actually unfolding, often catching them too early or too late to land inside the harm window. That matters more for real deployments than a headline AUC number, since a monitor that notices harm only after the agent has finished acting is not doing much safeguarding.
Most agent-safety evaluations still ask a simple yes-or-no question about whether a trajectory was harmful. This one suggests that framing has been letting monitors look better than they actually are at the one thing that counts in production: catching trouble while there is still time to stop it.