[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-rubricforge-judge-cuts-false-ai-agent-passes-by-a-third":10,"sections":45},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":35,"tags":36,"sources":40,"feedback":44,"feedback_at":22,"cost_usd":44,"total_tokens":44},5001,"rubricforge-judge-cuts-false-ai-agent-passes-by-a-third","RubricForge Judge Cuts False AI-Agent Passes by a Third","A new AI judge design cuts false passes in agent evaluation by about a third versus a standard rubric-based judge, though overall accuracy barely moves.","A new evaluation method builds an AI judge's scoring rubric from real pass\u002Ffail outcomes instead of guesswork, and it catches more of its own agent's failures.\n\nA new paper introduces RubricForge, a technique for building the rubric an AI judge uses to grade another AI agent's trajectories. Rather than hand-writing criteria, the common approach in tools like G-Eval, or fine-tuning the judge model itself, RubricForge evolves a rubric against a small set of trajectories with known outcomes, then freezes it and applies it to new trajectories in a single model call with no access to the underlying environment. The researchers tested it with one frozen 7B model acting as both agent and judge, on tau-bench (173 labeled trajectories from 220 rollouts) and WebShop (160 trajectories). Because the rubric is plain text, every verdict traces back to a named criterion rather than a black-box score.\n\nThe gain is not in raw accuracy. RubricForge's overall agreement with ground truth was not statistically distinguishable from a generic G-Eval judge, and its score calibration was slightly worse. What improved was the false-pass rate: on tau-bench, RubricForge wrongly credited a failed trajectory as a success 11.5 percent of the time, versus 17.3 percent for the generic judge, a meaningful cut but not the halving the framing might suggest. It also ranked graded WebShop outcomes more faithfully, with a Spearman correlation of 0.410 versus 0.370 for the baseline.\n\nThat distinction is the whole point: a false pass ships a broken agent, a false fail just costs a retry, so for anyone grading agents with agents, under-failing beats simply scoring well on paper.","[\"ai\",\"ai agents\",\"llm evaluation\",\"benchmarks\"]","2026-08-17T04:00:00.000Z","2026-08-17T04:51:50.532Z","2026-08-17T04:52:02.395Z","published",null,[24,30],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Fix the dek's claim that the method 'trades a sliver of raw accuracy' — the body states overall agreement with ground truth was NOT statistically different from the baseline judge, and only calibration (numeric score error) was slightly worse, so the dek should reflect a calibration trade-off, not an accuracy trade-off.","resolved",{"id":31,"reviewer":32,"round":33,"reason":34,"status":29},"publisher-r2","publisher",2,"The dek and body claim RubricForge 'halves' the false-pass rate, but the cited numbers (0.173 to 0.115) reflect only a ~33% reduction, not a halving — an internal inconsistency between the claimed and reported figures.","ai",[35,37,38,39],"ai agents","llm evaluation","benchmarks",[41],{"name":42,"url":43},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.13564",0,{"sections":46},[47,51,55,60,65,70,75,80,85,90,95,100,105,110],{"name":48,"slug":35,"count":49,"latest_published_at":50},"AI",3293,"2026-08-20T04:00:00.000Z",{"name":52,"slug":53,"count":54,"latest_published_at":50},"Security","security",435,{"name":56,"slug":57,"count":58,"latest_published_at":59},"Policy","policy",210,"2026-08-19T09:32:27.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":69},"Hardware","hardware",140,"2026-08-19T18:25:42.000Z",{"name":71,"slug":72,"count":73,"latest_published_at":74},"Consumer Tech","consumer-tech",95,"2026-08-18T16:05:00.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Science","science",90,"2026-08-19T18:41:02.000Z",{"name":81,"slug":82,"count":83,"latest_published_at":84},"Software","software",73,"2026-08-18T07:51:50.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":89},"Dev Tools","dev-tools",69,"2026-08-18T04:00:00.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"Startups","startups",47,"2026-08-19T19:13:46.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"Gaming","gaming",41,"2026-07-09T04:00:00.000Z",{"name":101,"slug":102,"count":103,"latest_published_at":104},"General","general",33,"2026-08-18T22:18:13.000Z",{"name":106,"slug":107,"count":108,"latest_published_at":109},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":111,"slug":112,"count":113,"latest_published_at":114},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]