AI/ ai · evaluation-bias · llm-benchmarks · research

Study Finds AI Judges Avoid Extreme Scores on Rating Scales

A new analysis of automated scoring models finds they systematically dodge the ends of a rating scale, and the bias can be fixed with targeted retraining.

AI models that auto-score text keep picking the safe middle option, even when they shouldn't.

Researchers tested four text-evaluation models: one they label JEV (version 1.13), and three open-weight systems they collectively call KEV. The paper uses these as internal labels rather than naming the actual products behind them. On the ANLI benchmark, JEV accurately classified 74.95% of examples, but still dumped 38.8% of all predictions into the Neutral category, accounting for 51.3% of its errors, despite gold labels and answer positions being balanced. Across 36 datasets with ordered rating scales, the four models used only 67% to 76% of the available range of answers, versus 87% to 102% on four unordered tasks. Expanding a scale from 2 points up to 14 points made things worse, not better: by 14 points, the models were using as little as 26% to 75% of the scale, even though their internal probability estimates stayed spread out.

That matters because these judge models are quietly replacing human raters in grading essays, scoring chatbot replies, and running sentiment analysis at scale, specifically because they are supposed to be fast and precise. If a model is handed a 1-to-10 scale and effectively treats it like a 1-to-4, every score built on top of that result is less meaningful than it looks. The researchers also showed the flaw is not baked in: a retraining method called BA-LoRA pushed scale usage from roughly 47% up to 86% on eight test scales.

It is a machine-learning version of a bias survey designers have fought for decades: people cluster toward the middle of any scale instead of using its extremes. Finding out that automated judges inherited the same flinch is a useful reality check for anyone treating them as a precise, unbiased stand-in for human judgment.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →