[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-fast-readouts-of-ai-judges-overstate-their-bias-study-finds":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},9587,"fast-readouts-of-ai-judges-overstate-their-bias-study-finds","Fast Readouts of AI Judges Overstate Their Bias, Study Finds","A new study finds reading an AI judge's first token overstates position bias in every test case, though it barely dents accuracy in most conditions.","Reading an AI judge's verdict off its first token is cheap, and it quietly inflates how biased the judge looks.\n\nA new study tested that shortcut, used by constrained-decoding and likelihood-scoring evaluation setups, against actually letting the judge finish its answer, across five models: three Qwen3 judges, Llama-3.1-8B, and Phi-3.5-mini. Those judges don't always lead with a verdict: on 12% to 49% of pairs for the Qwen3 models, and under 3% for Llama-3.1-8B and Phi-3.5-mini, the shortcut just returns whichever response was shown first instead of an actual judgment. Pooled across the 924 pairs where a judge hadn't committed to an answer yet, that forced read flipped 89.7% of the time when the two responses were swapped, versus 47.5% when the judge was allowed to finish generating.\n\nIn seven of the ten test conditions, that inflated bias score is mostly theater: it moves the measured position bias by 42 points while shifting actual judge accuracy by less than a point. The other three conditions didn't follow that pattern, so the shortcut isn't uniformly harmless, and anyone using it to judge a model's real-world reliability rather than just its optics needs to check which bucket applies. A smaller, separate glitch shows up even when judges do lead with a verdict token: on up to 5.5% of pairs, they open with one letter and then reason their way to the opposite answer.\n\nConstrained decoding and likelihood scoring are already standard in most evaluation harnesses, so this isn't a lab curiosity; it's baked into how a lot of judge benchmarks get produced right now.","[\"ai\",\"llm-judges\",\"benchmarking\",\"research\"]","2026-10-02T04:00:00.000Z","2026-10-03T01:52:45.716Z","2026-10-03T01:52:52.040Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The dek and body claim the first-token shortcut moves bias 'without changing accuracy' \u002F 'barely touching' it, but the source only supports that for seven of ten conditions — note the exception in the other three conditions instead of generalizing universally.","resolved","ai",[30,32,33,34],"llm-judges","benchmarking","research",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.00054",0,{"sections":41},[42,45,49,53,58,62,66,71,76,81,86,91,96,101],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",5896,{"name":46,"slug":47,"count":48,"latest_published_at":18},"Security","security",837,{"name":50,"slug":51,"count":52,"latest_published_at":18},"Policy","policy",438,{"name":54,"slug":55,"count":56,"latest_published_at":57},"Deals","deals",317,"2026-10-01T22:00:00.000Z",{"name":59,"slug":60,"count":61,"latest_published_at":18},"Hardware","hardware",199,{"name":63,"slug":64,"count":65,"latest_published_at":18},"Science","science",171,{"name":67,"slug":68,"count":69,"latest_published_at":70},"Consumer Tech","consumer-tech",155,"2026-10-01T19:54:10.000Z",{"name":72,"slug":73,"count":74,"latest_published_at":75},"Dev Tools","dev-tools",96,"2026-10-01T16:57:03.000Z",{"name":77,"slug":78,"count":79,"latest_published_at":80},"Software","software",93,"2026-09-30T21:41:11.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Startups","startups",90,"2026-10-01T21:55:22.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"Gaming","gaming",53,"2026-10-02T02:50:39.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":102,"slug":103,"count":104,"latest_published_at":105},"How-To","how-to",7,"2026-10-01T09:00:00.000Z"]