AI/ ai · llm-as-a-judge · ai-evaluation · research

New Paper Questions Whether AI Judges Actually Help

A new study finds that AI judges scoring high on accuracy don't necessarily make AI systems perform better, complicating how judges are used and trained.

Grading an AI judge on how accurate it is doesn't tell you whether it actually makes an AI system better.

That's the finding from a new paper on arXiv (arXiv:2609.37145, "From Judgment Quality to Downstream Utility: Rethinking LLM-as-a-Judge for Open-Ended Tasks", posted September 30, 2026). Researchers studied LLM-as-a-Judge systems, AI models used to score other AI models' answers on open-ended tasks that have no single right answer, like essay writing or advice giving. They varied how judges are set up along three axes: how fine-grained the verdicts are, whether judges write critiques alongside scores, and whether responses are evaluated one at a time or in batches. They then tested those judges not just as training-time reward signals, but as tools used during inference itself, through Best-of-N selection, judge-guided revision, and beam search.

The headline result: judgment quality and downstream usefulness don't reliably track each other. A judge that scores well on standard accuracy benchmarks can still fail to improve outcomes when plugged into training or generation, and the specific protocol used to elicit judgments, not just the underlying model, shapes both how good the judgments are and how much they help. Judge guidance did reliably convert extra test-time compute into performance gains, though how much varied by inference strategy.

That's a relevant wrinkle for an industry that has quietly outsourced a lot of evaluation, and even reward modeling, to another LLM grading the work: if the grader's report card doesn't predict real-world usefulness, a lot of RLHF pipelines and leaderboard rankings built on judge models are resting on an untested assumption.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →