AI models are increasingly the ones deciding whether other AI models did a good job, and it turns out they don't agree on what "good" means.
Researchers built a framework called JudgeProfile to figure out why. They tested 21 large language models acting as judges on 50,013 pairs of responses pulled from 17 public datasets, scoring each pair across 87 separate attributes like clarity, correctness, and level of detail. The surprising part: judges mostly agree on those individual attribute scores. Where they diverge is in how much weight they give each one when picking an overall winner. By estimating those weights and then reweighting them against reference labels, the researchers pushed agreement with reference labels from 66.48% up to 71.97%, beating both fine-tuning and rubric prompting.
That distinction matters because LLM judges have quietly become infrastructure. They grade chatbot outputs, rank responses for reinforcement learning, and populate leaderboards that companies cite as evidence their model is "better." If two judges disagree not because one perceives quality differently but because they prioritize differently, then swapping judges or writing better rubrics won't fix the inconsistency - you have to adjust what the judge values, not what it notices.
It's a tidier explanation than the usual hand-wave about bias, and a cheaper fix than retraining a judge model from scratch. Whether reweighting attributes holds up outside curated benchmark pairs is the next question worth asking.