[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-fast-ai-decision-models-underperform-and-so-did-the-studys-own-math":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},9860,"fast-ai-decision-models-underperform-and-so-did-the-studys-own-math","Fast AI Decision Models Underperform and So Did the Study's Own Math","A paired evaluation of two fast AI decision models finds one dramatically outperforms the other, and uncovers errors in its own earlier accuracy claims.","A new benchmark pits two fast AI decision models against each other, then checks its own math.\n\nResearchers evaluated two 'System-1' models built to make small, fast decisions inside AI agent systems, like picking which tool to use or flagging suspicious input, without calling a full language model each time. Across 11 decision tests built from 18 public datasets, the hosted model, Jev, beat the open-weight model, Laya, on 9 of them, by margins of 10.8 to 46 percentage points. Both models failed at routing tasks to the right underlying model, scoring no better than chance, and tied on judging whether retrieved documents were relevant. Laya also proved brittle: it flipped 30% of its answers when answer choices were simply reordered, and its accuracy collapsed to 31% when choosing among 50 similar tool options, compared with 98% for Jev on clearer cases.\n\nThe more useful finding might be the self-audit. Reviewing their own numbers, the researchers found three analysis errors and one design confound that had inflated earlier deployment claims. A missing pre-screening cost turned a reported 23.9% savings into an actual 4.3%. A gate-accuracy figure got reported as overall pipeline quality, conflating a 58% result with a 98% one. Error thresholds were tuned on the same data used to test them, missing held-out targets by up to 17%. Separately, a 'channel effect' made injection-detection false positives look worse than they were, an effect that vanished once tested with content native to the real deployment channel.\n\nBenchmarks that survive their authors' own fact-checking are rare. This one earns points for finding its own mistakes before a reviewer, or a deployment, did.","[\"ai agents\",\"llm evaluation\",\"ai research\",\"benchmarking\"]","2026-10-05T04:00:00.000Z","2026-10-05T10:53:41.476Z","2026-10-05T10:53:48.216Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The piece claims the researchers found 'three errors and one design flaw' but only describes two of the four (the cost-saving miscalculation and the gate-accuracy mix-up) — either enumerate all three analysis errors plus the design confound (the in-sample threshold miss and the channel-effect confound from the source) or drop the specific count.","resolved","ai",[32,33,34,35],"ai agents","llm evaluation","ai research","benchmarking",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.02267",0,{"sections":42},[43,46,50,55,60,65,69,74,78,82,87,92,97,102],{"name":44,"slug":30,"count":45,"latest_published_at":18},"AI",6167,{"name":47,"slug":48,"count":49,"latest_published_at":18},"Security","security",859,{"name":51,"slug":52,"count":53,"latest_published_at":54},"Policy","policy",444,"2026-10-03T15:02:01.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Deals","deals",323,"2026-10-04T13:00:00.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Hardware","hardware",204,"2026-10-03T14:50:50.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":18},"Science","science",177,{"name":70,"slug":71,"count":72,"latest_published_at":73},"Consumer Tech","consumer-tech",158,"2026-10-03T03:21:12.000Z",{"name":75,"slug":76,"count":77,"latest_published_at":18},"Dev Tools","dev-tools",97,{"name":79,"slug":80,"count":77,"latest_published_at":81},"Software","software","2026-10-04T10:00:00.000Z",{"name":83,"slug":84,"count":85,"latest_published_at":86},"Startups","startups",92,"2026-10-04T14:36:25.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"Gaming","gaming",53,"2026-10-02T02:50:39.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"General","general",51,"2026-10-05T02:35:01.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"Reviews","reviews",32,"2026-10-02T18:00:00.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"How-To","how-to",7,"2026-10-01T09:00:00.000Z"]