AI/ ai · llm-judges · research · benchmarks

Study Finds Hidden Signal Predicts When AI Judges Flip

Researchers show that an AI model's internal activations can predict when swapping the order of two answers would flip its verdict, beating simpler signals.

AI judges flip their verdicts depending on which answer comes first, and researchers can now predict when that will happen before it happens.

Researchers tested whether looking inside an LLM's residual stream, the internal activations recorded right before it issues a verdict, could flag cases where swapping the order of two candidate answers would change its judgment. They trained simple linear probes on 534 pairs from the JudgeBench benchmark, testing three Qwen3 judges and Llama-3.1-8B. The probes predicted order-sensitive flips with accuracy scores, measured in AUROC, between .621 and .850, beating a baseline that combined the model's stated confidence, verdict-label logits, response length, and its initial choice by up to .113 points. The same probes, trained once on JudgeBench and never recalibrated, still scored .685 to .853 on 1,802 unrelated comparisons from MT-Bench.

Checking both answer orders for every judgment doubles the compute cost of using an LLM as a grader, which is already standard practice for scoring chatbots, benchmarking models, and generating reinforcement learning signals. This method could flag only the unstable judgments for a second pass, instead of re-running everything. That matters because LLM-as-judge setups are quietly deciding leaderboard rankings and training rewards across the industry.

It is a patch, not a cure. The judges are still biased by candidate order; researchers just found a cheaper way to spot when the bias is about to strike.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →