A new benchmark says today's AI agents are far too eager to hit the brakes.
Researchers built SteerBench-Work, a 106-scenario test that puts AI agents at the moment right before they take a real action - sending an email, merging code, wiring money - and asks whether they should proceed or hold for a human. Each scenario is anchored to a real public incident and paired with an evidence-reversed mirror, so the same story can be tested with the risk flipped. Across 30 model conditions, agents wrongly blocked authorized, evidence-cleared work 28.1% of the time, compared to just 1.0% where they wrongly let something unsafe through. The gap widens on the reversed mirrors, where models scored 63.8% versus 98.5% on the original, well-known incidents.
That asymmetry matters because a workplace agent that refuses too often is a productivity tax, not a safety feature - someone still has to review and re-approve everything it stalls on. The researchers also found that more capable models often over-refuse more, not less, and that adding reasoning steps can patch a badly calibrated gate without fixing the underlying judgment.
In other words, the industry's go-to fix for AI reliability - throw a bigger or more 'reasoning' model at it - does not touch the real problem of knowing when to trust evidence that is already sitting in front of it.