A new paper offers a fix for a well-known flaw in AI evaluation: judges that favor whichever answer they see first.
The method, called DIAL, tackles a problem baked into how AI labs now grade model outputs: using other language models as automated judges. Prior research has shown these LLM judges can flip their verdict simply based on response order, and even after correcting for that, their rankings still drift from what human evaluators would actually prefer. DIAL combines a large volume of cheap LLM-generated comparisons with a small set of human-labeled ones to separate out judge-specific position effects from genuine preference signal, then calibrates the result toward human judgment. The researchers back this with theoretical work on identifying the bias and quantifying uncertainty in the corrected estimates, and they tested it on three human-preference benchmarks plus a new dataset of over 410,000 judgments from 21 LLM judges evaluated in both response orders.
LLM-as-judge has become the default way labs and benchmark makers grade model outputs, because paying humans to rate everything doesn't scale. But if the judge itself is skewed by a formatting quirk like answer order, every leaderboard and reward signal built on it inherits that skew. DIAL's pitch is a way to correct for that using far fewer human labels than a full re-annotation would require.
It's still an unreviewed preprint, and the fix only helps if labs adopt a new calibration step instead of the cheaper workaround most already lean on: swapping answer order and averaging the results.