AI/ llm-as-a-judge · ai-evaluation · benchmarking · research

AI Judges Are Mostly Reliable, Until You Change One Setting

A new study finds AI judge models are accurate on average, but swapping which model does the judging can swing leniency by up to 56 percentage points.

Researchers just showed that AI models used to grade other AI models' work are trustworthy on average, but surprisingly easy to tilt with a small tweak to their setup.

A new arXiv paper tested 10 reasoning models acting as "LLM-as-a-judge" evaluators across two tasks: rating sentence sentiment and toxicity on a 1-7 scale (over 500 items per category), and judging whether question-answer pairs were correct (600 pairs). On the rating task, judges missed human scores by an average of just 0.11 points out of 7, and the toxicity judges actually beat standard classifiers. On the accuracy task, judges got it right 96.5% of the time on average. But swapping the rating scale alone shifted measured bias by up to 0.93 points, and switching which model did the judging moved leniency - how often it called an answer correct - by as much as 56.1 percentage points. A more detailed prompt dropped leniency by 28.9 points. Oddly, dialing down how hard the model "thinks" before answering changed neither accuracy nor leniency.

LLM judges have quietly become the default grading rubric for a huge share of AI research and product evaluation, replacing slower, pricier human review. This study says the judges are fine in aggregate, but the specific setup - which model, which scale, how verbose the prompt - can move the result more than the thing being tested actually changes.

So before trusting a benchmark score, it is worth asking who, or what, graded it.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →