[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-benchmark-exposes-how-ai-models-fail-at-legal-judgment":10,"sections":39},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":34,"feedback":38,"feedback_at":22,"cost_usd":38,"total_tokens":38},9460,"new-benchmark-exposes-how-ai-models-fail-at-legal-judgment","New Benchmark Exposes How AI Models Fail at Legal Judgment","JusticeAxis tests AI on 256 real criminal cases and finds scale shifts how models fail, though a lightweight fix brings open models up to commercial standard.","AI models tested on real criminal cases fail in two opposite ways, and simply making them bigger does not fix it.\n\nResearchers released JusticeAxis, a benchmark built from 256 real criminal cases across 18 countries, each including audio, image, and text evidence. For every case, lawyers wrote the actual judgment plus two deliberately flawed versions representing specific ways judgment can go wrong. Testing showed smaller open-weight models tend to invent justifications not supported by the case record, while larger frontier models tend to apply the strict letter of the law and ignore context that should change the outcome. The researchers also built JusticeAgent, a system that separates fact-finding from judgment and adds 'skills' learned from past case outcomes, filtered through a statistical confidence check before they're trusted.\n\nThat split matters because it reframes 'better AI judgment' as two distinct failure modes, not one quality dial that scale slides up. The more striking result: bolting JusticeAgent onto a frozen open-weight model, with no retraining, brought its performance up to the level of commercial frontier models.\n\nIt's one paper's benchmark, not a courtroom deployment, and 'commercial level' performance on a research dataset is a long way from being trusted with an actual sentence.","[\"ai\",\"legal-tech\",\"benchmarks\"]","2026-10-02T04:00:00.000Z","2026-10-02T20:11:05.322Z","2026-10-02T20:11:10.438Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Add the paper's other headline result — that JusticeAgent raised a frozen open-weight model to commercial-level performance — since omitting it makes the closing 'scale didn't fix anything' takeaway misleading about what the source actually found.","resolved","ai",[30,32,33],"legal-tech","benchmarks",[35],{"name":36,"url":37},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.00353",0,{"sections":40},[41,44,48,53,58,63,68,73,78,83,88,93,98,103],{"name":42,"slug":30,"count":43,"latest_published_at":18},"AI",5765,{"name":45,"slug":46,"count":47,"latest_published_at":18},"Security","security",831,{"name":49,"slug":50,"count":51,"latest_published_at":52},"Policy","policy",437,"2026-10-01T18:10:00.000Z",{"name":54,"slug":55,"count":56,"latest_published_at":57},"Deals","deals",317,"2026-10-01T22:00:00.000Z",{"name":59,"slug":60,"count":61,"latest_published_at":62},"Hardware","hardware",198,"2026-10-01T17:38:48.000Z",{"name":64,"slug":65,"count":66,"latest_published_at":67},"Science","science",168,"2026-10-01T18:35:55.000Z",{"name":69,"slug":70,"count":71,"latest_published_at":72},"Consumer Tech","consumer-tech",155,"2026-10-01T19:54:10.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":77},"Dev Tools","dev-tools",96,"2026-10-01T16:57:03.000Z",{"name":79,"slug":80,"count":81,"latest_published_at":82},"Software","software",93,"2026-09-30T21:41:11.000Z",{"name":84,"slug":85,"count":86,"latest_published_at":87},"Startups","startups",90,"2026-10-01T21:55:22.000Z",{"name":89,"slug":90,"count":91,"latest_published_at":92},"Gaming","gaming",53,"2026-10-02T02:50:39.000Z",{"name":94,"slug":95,"count":96,"latest_published_at":97},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":99,"slug":100,"count":101,"latest_published_at":102},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":104,"slug":105,"count":106,"latest_published_at":107},"How-To","how-to",7,"2026-10-01T09:00:00.000Z"]