A new study shows that the tests used to judge AI debugging tools can accidentally hand over the answer before any real diagnosis happens.
Researchers testing an automated debugger for closed-loop decision agents found that two different fault-finding methods, an exact "minimum hitting set" solver and a faster greedy approximation, produced identical results in all 12 development test cases, and matched the known injected fault in 9 of those 12. An audit found why: the verifier's built-in checks, called exact-anchor predicates, were directly producing the planted fault pair in all 9 cases, meaning the test already contained the answer. After stripping out those shortcut checks, the recovery rate dropped to 8 of 9 cases. The team then built a stricter evaluation process requiring evidence to clear two gates, repeated exposure and matched runtime evidence, before any detection algorithm gets credit for a result.
This is a quiet but important warning for anyone building or benchmarking AI agents: a test that looks like it measures algorithm quality can actually just be grading its own design. Teams comparing agent tools, prompts, or policies on leaderboard-style benchmarks should ask whether the setup could be leaking labels, not just whether the scores look good. In a new heldout evaluation spanning 1,440 cases, the stricter method still held up, with a false-admission rate bounded below roughly 14 percent at 95 percent confidence.
It is the AI-era version of teaching to the test, except here the test was unknowingly writing itself the answer key.