Reading an AI judge's verdict off its first token is cheap, and it quietly inflates how biased the judge looks.
A new study tested that shortcut, used by constrained-decoding and likelihood-scoring evaluation setups, against actually letting the judge finish its answer, across five models: three Qwen3 judges, Llama-3.1-8B, and Phi-3.5-mini. Those judges don't always lead with a verdict: on 12% to 49% of pairs for the Qwen3 models, and under 3% for Llama-3.1-8B and Phi-3.5-mini, the shortcut just returns whichever response was shown first instead of an actual judgment. Pooled across the 924 pairs where a judge hadn't committed to an answer yet, that forced read flipped 89.7% of the time when the two responses were swapped, versus 47.5% when the judge was allowed to finish generating.
In seven of the ten test conditions, that inflated bias score is mostly theater: it moves the measured position bias by 42 points while shifting actual judge accuracy by less than a point. The other three conditions didn't follow that pattern, so the shortcut isn't uniformly harmless, and anyone using it to judge a model's real-world reliability rather than just its optics needs to check which bucket applies. A smaller, separate glitch shows up even when judges do lead with a verdict token: on up to 5.5% of pairs, they open with one letter and then reason their way to the opposite answer.
Constrained decoding and likelihood scoring are already standard in most evaluation harnesses, so this isn't a lab curiosity; it's baked into how a lot of judge benchmarks get produced right now.