AI/ llm-judges · ai-evaluation · ai-research · benchmarking

LLM Judge Panels Share Errors, Undermining Consensus

A new study shows LLM judges often fail on the same examples, meaning judge panels give far less independent confirmation than their vote counts suggest.

A new study finds that stacking LLM judges together does not multiply your confidence the way most benchmarks assume.

Researchers tested ten LLM judges, spanning open-weight and frontier models, and measured how correlated their mistakes were. They found an average pairwise error correlation of 0.21, meaning judges tend to miss the same cases rather than fail independently. That correlation gets stronger, not weaker, among the highest-accuracy frontier judges, even ones built by different companies. The practical effect: those ten judges carry roughly the statistical weight of 3.5 independent judges, not ten.

That gap matters because panels are routinely used to decide which AI system performed better. The study found that in up to 28% of comparisons, ignoring shared errors makes one system look meaningfully better, a conclusion that disappears once the correlation is accounted for. The researchers also note that how errors are distributed, spread across most judges or clustered in a subset, changes which voting method actually works.

Their proposed fix is refreshingly low-tech: check judges against a small set of trusted examples first, map out where they share mistakes, and pick a voting method based on that before trusting the panel's verdict on anything new.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →