[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-open-source-ai-judge-model-barely-beats-a-coin-flip":10,"sections":35},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":30,"feedback":34,"feedback_at":22,"cost_usd":34,"total_tokens":34},10632,"open-source-ai-judge-model-barely-beats-a-coin-flip","Open-Source AI Judge Model Barely Beats a Coin Flip","A preregistered benchmark test found open model CLM-v0.1-8B judging near chance, while similarly sized reward and generative judges scored far higher.","A fully open 8-billion-parameter model built to judge other AI systems' outputs scores about as well as a coin flip.\n\nResearchers tested Contrastive-LM's CLM-v0.1-8B on five public preference benchmarks plus one hallucination benchmark, under evaluation rules locked in before anyone saw a test item. The model scored between 0.351 and 0.593 on tasks where blind guessing already clears 0.25 to 0.5, and it was statistically indistinguishable from chance on RM-Bench and JudgeBench. On the hallucination benchmark, HaluEval, it returned the same label on every item, matching a baseline that just always picks the first answer. A reward model and a generative judge with the same parameter count, tested under identical conditions, scored 0.764-0.976 and 0.611-0.778, and every gap was statistically significant.\n\nThe one bright spot is calibration. CLM's raw confidence scores run overconfident by as much as 0.401, but a single temperature adjustment fit on held-out data brings that error down to 0.062, and the fixed confidence can flag the model's own mistakes better than chance on three of six benchmarks. That sounds like a foundation for a cheap-first, escalate-when-unsure pipeline. It isn't much of one: routing low-confidence cases to a stronger judge still had to hand off 92.3 to 100 percent of items to hit the preregistered accuracy bar.\n\nKnowing exactly how wrong a model is, it turns out, is not the same as the model being right.","[\"llm-as-judge\",\"ai benchmarks\",\"open-source ai\",\"calibration\"]","2026-10-07T04:00:00.000Z","2026-10-09T00:14:07.207Z","2026-10-09T00:14:10.856Z","published",null,[],"ai",[26,27,28,29],"llm-as-judge","ai benchmarks","open-source ai","calibration",[31],{"name":32,"url":33},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.07177",0,{"sections":36},[37,41,46,51,56,61,65,70,75,79,84,89,94,99],{"name":38,"slug":24,"count":39,"latest_published_at":40},"AI",6506,"2026-10-07T18:45:00.000Z",{"name":42,"slug":43,"count":44,"latest_published_at":45},"Security","security",911,"2026-10-07T19:53:42.000Z",{"name":47,"slug":48,"count":49,"latest_published_at":50},"Policy","policy",474,"2026-10-07T18:23:21.000Z",{"name":52,"slug":53,"count":54,"latest_published_at":55},"Deals","deals",453,"2026-10-07T23:58:31.000Z",{"name":57,"slug":58,"count":59,"latest_published_at":60},"Hardware","hardware",222,"2026-10-07T21:19:54.000Z",{"name":62,"slug":63,"count":64,"latest_published_at":18},"Science","science",187,{"name":66,"slug":67,"count":68,"latest_published_at":69},"Consumer Tech","consumer-tech",174,"2026-10-07T17:41:41.000Z",{"name":71,"slug":72,"count":73,"latest_published_at":74},"Software","software",113,"2026-10-07T18:10:00.000Z",{"name":76,"slug":77,"count":73,"latest_published_at":78},"Startups","startups","2026-10-07T23:36:57.000Z",{"name":80,"slug":81,"count":82,"latest_published_at":83},"Dev Tools","dev-tools",105,"2026-10-07T16:59:11.000Z",{"name":85,"slug":86,"count":87,"latest_published_at":88},"General","general",61,"2026-10-07T22:00:24.000Z",{"name":90,"slug":91,"count":92,"latest_published_at":93},"Gaming","gaming",56,"2026-10-07T12:00:00.000Z",{"name":95,"slug":96,"count":97,"latest_published_at":98},"Reviews","reviews",33,"2026-10-05T11:57:17.000Z",{"name":100,"slug":101,"count":102,"latest_published_at":103},"How-To","how-to",8,"2026-10-05T09:00:00.000Z"]