[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-benchmark-shows-ai-error-checkers-fail-outside-math-problems":10,"sections":34},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":29,"feedback":33,"feedback_at":22,"cost_usd":33,"total_tokens":33},9225,"benchmark-shows-ai-error-checkers-fail-outside-math-problems","Benchmark shows AI error checkers fail outside math problems","A new benchmark called LSR-Ben finds that process reward models and LLMs are far worse at catching reasoning errors in science and logic than in math.","A new benchmark finds that the AI models meant to catch mistakes in step-by-step reasoning are good at it mostly when the problem is math.\n\nResearchers built LSR-Ben, a benchmark that tests process reward models (PRMs) - systems that review a model's intermediate reasoning steps for errors - on scientific and logical reasoning tasks spanning nine subdomains, rather than the math problems most existing tests rely on. They ran 22 models, a mix of dedicated PRMs and general-purpose LLMs, through the benchmark. Error detection dropped off sharply once the material moved beyond math. The team also found a split in checking style: LLMs tend to flag too many steps as wrong, while PRMs tend to miss errors that are actually there.\n\nProcess reward models are supposed to be the component that makes test-time scaling work, catching a bad reasoning step before it compounds into a wrong final answer. If that oversight layer only holds up on math, it undercuts claims that these systems generalize to the reasoning tasks people actually want AI for, like scientific analysis or logical deduction.\n\nMost reasoning benchmarks to date have leaned on math because it's easy to grade automatically. LSR-Ben's finding - that today's error checkers are narrower than advertised - says as much about the gaps in how we test these models as it does about the models themselves.","[\"ai\",\"benchmarks\",\"llms\",\"reasoning\"]","2026-10-01T04:00:00.000Z","2026-10-02T02:49:09.693Z","2026-10-02T02:49:15.151Z","published",null,[],"ai",[24,26,27,28],"benchmarks","llms","reasoning",[30],{"name":31,"url":32},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2605.01203",0,{"sections":35},[36,39,43,47,52,57,61,66,71,75,80,85,90,95],{"name":37,"slug":24,"count":38,"latest_published_at":18},"AI",5629,{"name":40,"slug":41,"count":42,"latest_published_at":18},"Security","security",816,{"name":44,"slug":45,"count":46,"latest_published_at":18},"Policy","policy",430,{"name":48,"slug":49,"count":50,"latest_published_at":51},"Deals","deals",298,"2026-09-30T21:00:26.000Z",{"name":53,"slug":54,"count":55,"latest_published_at":56},"Hardware","hardware",196,"2026-09-30T13:00:00.000Z",{"name":58,"slug":59,"count":60,"latest_published_at":18},"Science","science",163,{"name":62,"slug":63,"count":64,"latest_published_at":65},"Consumer Tech","consumer-tech",149,"2026-09-30T22:57:11.000Z",{"name":67,"slug":68,"count":69,"latest_published_at":70},"Dev Tools","dev-tools",93,"2026-10-01T02:30:48.000Z",{"name":72,"slug":73,"count":69,"latest_published_at":74},"Software","software","2026-09-30T21:41:11.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Startups","startups",84,"2026-09-30T20:39:09.000Z",{"name":81,"slug":82,"count":83,"latest_published_at":84},"Gaming","gaming",51,"2026-09-30T16:24:30.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":89},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]