A fully open 8-billion-parameter model built to judge other AI systems' outputs scores about as well as a coin flip.
Researchers tested Contrastive-LM's CLM-v0.1-8B on five public preference benchmarks plus one hallucination benchmark, under evaluation rules locked in before anyone saw a test item. The model scored between 0.351 and 0.593 on tasks where blind guessing already clears 0.25 to 0.5, and it was statistically indistinguishable from chance on RM-Bench and JudgeBench. On the hallucination benchmark, HaluEval, it returned the same label on every item, matching a baseline that just always picks the first answer. A reward model and a generative judge with the same parameter count, tested under identical conditions, scored 0.764-0.976 and 0.611-0.778, and every gap was statistically significant.
The one bright spot is calibration. CLM's raw confidence scores run overconfident by as much as 0.401, but a single temperature adjustment fit on held-out data brings that error down to 0.062, and the fixed confidence can flag the model's own mistakes better than chance on three of six benchmarks. That sounds like a foundation for a cheap-first, escalate-when-unsure pipeline. It isn't much of one: routing low-confidence cases to a stronger judge still had to hand off 92.3 to 100 percent of items to hit the preregistered accuracy bar.
Knowing exactly how wrong a model is, it turns out, is not the same as the model being right.