[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-study-finds-ai-code-judges-often-just-call-it-a-tie":10,"sections":49},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":39,"tags":40,"sources":44,"feedback":48,"feedback_at":22,"cost_usd":48,"total_tokens":48},8172,"study-finds-ai-code-judges-often-just-call-it-a-tie","Study Finds AI Code Judges Often Just Call It a Tie","A label-free diagnostic shows why AI systems that judge code often call comparisons a tie, and letting them abstain recovers real accuracy.","When an AI model grades another AI's code, it often just shrugs and calls it a draw.\n\nResearchers tested MARCH, a published multi-agent code-judging framework that breaks a verdict into checkable claims, across 80 condition-by-cell measurements on two code-judging benchmarks. The system declared both solutions in a comparison equally good on 78 to 95 percent of cases. Counting those tie-outs, its overall accuracy at picking the better solution came to 4.4 percent, worse than simply asking the same underlying model to make a direct call, which got it right 43.7 percent of the time. Making the problems easier, or using a bigger judge model, did not change the result.\n\nThe researchers trace the failure to a gap in how multi-agent verification works: it checks claims against evidence, and that works well when the evidence is retrieved documents, because those are independent of the answer and differ between the two things being compared. Code does not reliably meet either condition, so the judge has nothing solid left to break ties with. Two measurements pulled from the pipeline's own logs, no ground-truth labels required, can flag exactly when that is happening. Using one of them as a gate to let the judge decline comparisons it cannot support lifts accuracy on the comparisons it still answers, about half the total, from a 20.7 percent baseline on that same answerable subset to 36.9 percent.\n\nThe fix here is not a smarter judge. It is a judge willing to say \"I don't know,\" which is a lower bar than it sounds, and one a surprising amount of AI tooling still fails to clear.","[\"ai\",\"multi-agent-systems\",\"code-review\",\"llm-evaluation\"]","2026-09-28T04:00:00.000Z","2026-09-28T12:40:22.586Z","2026-09-28T12:40:29.017Z","published",null,[24,30,35],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The source only describes MARCH as 'a published multi-agent verification framework,' not adopted or well-known — remove or substantiate the dek\u002Flede's characterization of it as 'popular' since that's an unsupported claim about its adoption.","resolved",{"id":31,"reviewer":32,"round":33,"reason":34,"status":29},"publisher-r2","publisher",2,"The body cites MARCH's baseline accuracy as 4.4 percent early on but then says gating 'raised accuracy from 20.7 to 36.9 percent' without explaining the discrepancy between the two baseline figures.",{"id":36,"reviewer":32,"round":37,"reason":38,"status":29},"publisher-r3",3,"The accuracy figures are inconsistent — MARCH's overall accuracy is stated as 4.4 percent, but the abstention-based version is later said to rise 'from 20.7 percent to 36.9 percent' without explaining how 20.7 percent relates to the earlier 4.4 percent figure.","ai",[39,41,42,43],"multi-agent-systems","code-review","llm-evaluation",[45],{"name":46,"url":47},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.30328",0,{"sections":50},[51,54,58,63,68,73,77,82,87,92,97,102,106,111],{"name":52,"slug":39,"count":53,"latest_published_at":18},"AI",4796,{"name":55,"slug":56,"count":57,"latest_published_at":18},"Security","security",762,{"name":59,"slug":60,"count":61,"latest_published_at":62},"Policy","policy",399,"2026-09-27T18:39:02.000Z",{"name":64,"slug":65,"count":66,"latest_published_at":67},"Deals","deals",261,"2026-09-27T15:30:35.000Z",{"name":69,"slug":70,"count":71,"latest_published_at":72},"Hardware","hardware",188,"2026-09-27T20:46:36.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":18},"Science","science",151,{"name":78,"slug":79,"count":80,"latest_published_at":81},"Consumer Tech","consumer-tech",135,"2026-09-26T14:30:00.000Z",{"name":83,"slug":84,"count":85,"latest_published_at":86},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"Dev Tools","dev-tools",84,"2026-09-26T04:20:58.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"Startups","startups",76,"2026-09-25T18:33:59.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":103,"slug":104,"count":100,"latest_published_at":105},"General","general","2026-09-26T17:02:42.000Z",{"name":107,"slug":108,"count":109,"latest_published_at":110},"Reviews","reviews",30,"2026-09-24T20:07:31.000Z",{"name":112,"slug":113,"count":114,"latest_published_at":115},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]