A new paper audits the validator suite guarding a deployed generative AI agent, and finds most of the checks don't actually separate working output from broken output.
Researchers tested 13 validators against 550 runtime builds and 350 static builds, each labeled by whether it worked once shipped. Using a separation statistic (true positive rate minus false positive rate), only two checks held up after correcting for multiple comparisons, two more were weak, and nine were statistically indistinguishable from doing nothing - three of them never fired at all. The bigger problem: validators are more likely to get skipped on broken builds than on good ones. Across four runs covering 1,867 builds, probes were skipped on roughly 15.6 to 16.6 percent of broken builds but at most 0.3 percent of acceptable ones, and every skip was logged as a pass.
That skip-as-pass logging bug sets a hard ceiling: no matter how good a check is, if it depends on a live artifact, it can't catch more than about 84 percent of broken builds in this setup. The same blind spot shows up a layer above the individual checks, too. Across tens of thousands of judge-scored builds, 32.5 percent of rejections came with no recorded reason at all - a decision with no evidence trail.
The findings matter less for the body count of flawed checks and more for what they say about how deployment pipelines keep score. If a system that fails to run a test records that test as "passed," every uptime and reliability number built on top of it is inflated by construction. It's a familiar failure mode from ordinary software testing - a skipped CI job isn't a green checkmark - just showing up with higher stakes in an AI agent's release gate.