A safety check built to stop AI agents from taking harmful actions can be broken by editing a single line in the diagram it trusts.
Researchers tested a verifier called CIVeX, which gates an AI agent's tool calls - the actions that actually change something in the world - by checking whether a proposed action can be causally traced through a committed graph of how actions and states relate, then issuing a certificate with a confidence bound. On a benchmark built to include confounding (hidden factors that muddy cause and effect), CIVeX reported zero false executions. The researchers then red-teamed it by editing only that graph, not the verifier's logic. Deleting a single bidirected edge pushed false executions to 15.3%, with 91% of those harmful, and the agent's utility score fell from +2.27 to +0.35; flipping one arrowhead, so a genuine in-between cause got mislabeled as a confounder, produced 48.9% false executions and not a single correct one - and every one of those bad calls still carried a certificate that looked internally valid.
This matters because the certificate is the whole point of using a verifier like this: it's supposed to replace trusting the model's word with something you can check. The researchers' fix - testing each certified action against a small randomized sample before letting it run - caught both attacks, with just 2 false alarms across 555 executions on an honest graph. But under the same corrupted graph, that fix left 97.1% of genuinely beneficial actions unexecuted, because the bad graph had already rejected them before the check got a chance to run. Making the system safe again cost 127 extra test experiments per 1,050 actions; recovering the lost value on top of that cost 614 more - at which point the verifier was just making the same calls an honestly specified graph would have made, while spending its entire experiment budget to get there.
It's a reminder that "verified safe" is only as good as the spec it's verified against - the agent didn't get smarter here, it just got a lot more expensive to run.