When an AI model grades another AI's code, it often just shrugs and calls it a draw.
Researchers tested MARCH, a published multi-agent code-judging framework that breaks a verdict into checkable claims, across 80 condition-by-cell measurements on two code-judging benchmarks. The system declared both solutions in a comparison equally good on 78 to 95 percent of cases. Counting those tie-outs, its overall accuracy at picking the better solution came to 4.4 percent, worse than simply asking the same underlying model to make a direct call, which got it right 43.7 percent of the time. Making the problems easier, or using a bigger judge model, did not change the result.
The researchers trace the failure to a gap in how multi-agent verification works: it checks claims against evidence, and that works well when the evidence is retrieved documents, because those are independent of the answer and differ between the two things being compared. Code does not reliably meet either condition, so the judge has nothing solid left to break ties with. Two measurements pulled from the pipeline's own logs, no ground-truth labels required, can flag exactly when that is happening. Using one of them as a gate to let the judge decline comparisons it cannot support lifts accuracy on the comparisons it still answers, about half the total, from a 20.7 percent baseline on that same answerable subset to 36.9 percent.
The fix here is not a smarter judge. It is a judge willing to say "I don't know," which is a lower bar than it sounds, and one a surprising amount of AI tooling still fails to clear.