AI/ llm-as-judge · ai benchmarks · open-source ai · calibration

Open-Source AI Judge Model Barely Beats a Coin Flip

A preregistered benchmark test found open model CLM-v0.1-8B judging near chance, while similarly sized reward and generative judges scored far higher.

A fully open 8-billion-parameter model built to judge other AI systems' outputs scores about as well as a coin flip.

Researchers tested Contrastive-LM's CLM-v0.1-8B on five public preference benchmarks plus one hallucination benchmark, under evaluation rules locked in before anyone saw a test item. The model scored between 0.351 and 0.593 on tasks where blind guessing already clears 0.25 to 0.5, and it was statistically indistinguishable from chance on RM-Bench and JudgeBench. On the hallucination benchmark, HaluEval, it returned the same label on every item, matching a baseline that just always picks the first answer. A reward model and a generative judge with the same parameter count, tested under identical conditions, scored 0.764-0.976 and 0.611-0.778, and every gap was statistically significant.

The one bright spot is calibration. CLM's raw confidence scores run overconfident by as much as 0.401, but a single temperature adjustment fit on held-out data brings that error down to 0.062, and the fixed confidence can flag the model's own mistakes better than chance on three of six benchmarks. That sounds like a foundation for a cheap-first, escalate-when-unsure pipeline. It isn't much of one: routing low-confidence cases to a stronger judge still had to hand off 92.3 to 100 percent of items to hit the preregistered accuracy bar.

Knowing exactly how wrong a model is, it turns out, is not the same as the model being right.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →