A cheap AI judge can grade other AI's answers almost as well as GPT-6, as long as it knows when to ask for help.
Researchers built JEV, a judge model that skips writing out reasoning and instead returns probability scores for each possible verdict. Tested against sixteen other AI judges and checked against blinded human raters, JEV landed within three points of GPT-6's accuracy on judgments readable straight off the text, while charging 0.36% of GPT-6's fee and answering in a 0.15-second median. It fell behind on tasks that require actual derivation, like math, code, and logic problems. The researchers used JEV's own confidence score to flag those weak spots, then built a system that hands off only the uncertain cases to a full reasoning judge.
On 1,610 held-out test pairs, that hybrid setup beat GPT-6's accuracy by 0.9 points while costing 41% of what pure GPT-6 judging costs, and in a live test on two new workloads it matched GPT-6's accuracy exactly. That's a concrete answer to the complaint that's trailed LLM-as-a-judge since people started grading model outputs with other models: it's slow and expensive at scale.
The catch: the confidence trick gets shakier on style-adversarial phrasing and prose without a reference to check against, so this is less a drop-in replacement than a judgment call about which parts of an evaluation pipeline you're willing to hand to a fast guesser.