Letting an AI agent click the button isn't the hard part. Knowing when it should click is.
Researchers tested ten combinations of AI models and agent harnesses on SafeActBench, a new benchmark of 656 tasks spanning six operational domains and five task protocols. The benchmark does not just check whether an agent's final action was correct. It checks whether the agent had actually gathered enough proof before acting. The team built what they call a provenance-bound Evidence Ledger to log exactly what information an agent had confirmed and when, plus a deterministic evaluator that checks whether multi-step tasks hit their prerequisites in the right order. Agents were good at judging, in the abstract, whether an action was justified. They were far less reliable at following through on that judgment once execution started.
Most agent failures get blamed on models not knowing enough. This study locates the problem somewhere else: agents frequently stop investigating too early, or act before confirming evidence they themselves flagged as necessary. Once that evidence is actually in hand, single actions tend to go fine. The trouble compounds on multi-action workflows, where skipped prerequisites and half-finished steps stack up.
A chatbot that guesses wrong just embarrasses itself. An agent that guesses wrong deletes a file or fires off a transaction - which is why being well-reasoned and being well-executed turned out to be two separate scorecards.