AI/ ai-evaluation · llm-as-a-judge · benchmarks · research

Study Finds AI Judges Flip Verdicts When Order Is Swapped

A new study puts six AI models through repeated judging trials and finds their verdicts often flip based on answer order and repetition alone.

Six leading AI models were asked to grade other models' answers again and again - and they often disagreed with themselves.

Researchers ran six frontier language models through four benchmarks, testing five prompt formats, two answer orderings, and three temperature settings, with ten repeats per condition. Even at temperature zero, where outputs are supposed to be fixed, the same judge sometimes reached different verdicts on identical repeats of the same question. Swapping the order in which two answers were shown flipped the majority of verdicts on harder tasks. One model came out looking like the most consistent judge in the study simply because it repeated the same verdict every time - and that verdict matched the correct answer only 51% of the time, barely better than a coin flip.

The AI industry treats LLM judges as close to ground truth for scoring chatbots, benchmarks, and leaderboards. This research shows consistency and accuracy are not the same thing, and a judge can max out on one while quietly failing the other. The authors propose a new metric, the trustworthy verdict rate, that checks whether a verdict is reproducible, order-invariant, and correct all at once, rather than trusting any single pass.

If a leaderboard's ranking depends on which answer happened to go first, the fix may not be a better judge model at all - the study finds that switching from pairwise comparisons to holistic rubric scoring does more for reliability than any prompting trick.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →