Most AI safety guardrails watch for agents doing something forbidden. A new paper argues that misses the bigger problem: agents quietly skipping steps they were supposed to take.
Researchers studied trajectories from GLM-5.3 on a standard agent-safety benchmark and found 56.92% contained at least one "unfulfilled obligation" - a required safety action the agent never performed - compared with just 30.00% that contained an outright forbidden action. To test whether existing guard models can even catch these omissions, the team built ObligationBench, a 240-trajectory benchmark covering issue resolution, feature development, and terminal operations, with each case checked by human experts. Across 14 current guard models, the best recall was 48.97% and the best exact-match rate was just 10.00%. The researchers then trained their own model, ObligationGuard, on 40,000 synthetic examples, pushing those numbers to 57.52% recall and 21.67% exact match.
That reframes what "safe" means for an AI agent. An agent that never does anything forbidden can still be unsafe if it skips a required check, confirmation, or cleanup step, and today's guard models are built almost entirely to flag bad actions rather than notice absence. That is a structural blind spot in how agents get audited before being trusted with real tasks.
Even the paper's own fix catches barely a fifth of the omissions it goes looking for, which says less about the researchers and more about how early we still are in defining what an AI agent actually owes us.