Researchers say the entire LLM leaderboard industry has been solving the wrong problem.
A new paper, "Accounting for Bias Enables Sustainable LLM Evaluation" (arXiv:2609.31184), argues that LLM-as-a-judge evaluation, now the default way to rank chatbots and models, has a measurement problem, not a data problem. Judges carry documented biases: position bias, verbosity bias, inconsistent severity between judges, and a tendency to favor outputs from their own model family. Current leaderboards try to average out that noise by running more and more pairwise comparisons. The authors instead propose a latent variable model that jointly fits pairwise and ordinal judgments while explicitly correcting for those biases, producing reliable rankings from far fewer comparisons.
This matters because almost every model ranking in use today, from public chatbot arenas to internal eval pipelines, leans on LLM judges without accounting for the fact that those judges are not neutral instruments. If bias correction can substitute for brute-force scale, evaluation gets both cheaper and more defensible, since the fix targets the actual source of error instead of trying to bury it under volume. The paper also notes the correction model itself costs little compute compared to a single round of LLM inference, so this is not a tradeoff between rigor and sustainability.
It is a modest paper with an awkward implication: some share of the compute burned through leaderboard season may have been buying statistical reassurance, not accuracy.