[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-benchmark-catches-ai-models-faking-scientific-reasoning":10,"sections":45},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":35,"tags":36,"sources":40,"feedback":44,"feedback_at":22,"cost_usd":44,"total_tokens":44},7581,"new-benchmark-catches-ai-models-faking-scientific-reasoning","New Benchmark Catches AI Models Faking Scientific Reasoning","Sci-MMR shows leading multimodal AI often gets the right answer without ever assembling or reasoning through the actual evidence.","A new benchmark finds that top AI research assistants often land on the right scientific answer while skipping the evidence trail that's supposed to get them there.\n\nResearchers built Sci-MMR, a benchmark of 235 multi-step reasoning tasks spanning four scientific disciplines, each requiring models to pull evidence from an average of nine figure panels. Rather than just checking whether the final answer is correct, the benchmark traces whether that answer is actually backed by a full chain of evidence - citations, figures, and specific highlighted regions. Testing eight frontier multimodal models, the researchers found answer accuracy beat complete-evidence recovery by more than 20 percentage points. In plain terms: the models guess right more often than they can show their work.\n\nThe gap splits into two main failure types, plus a smaller leftover the study doesn't break down further. About 57.2% of errors trace to models failing to extract complete evidence from scientific figures in the first place - handing them the correct evidence directly boosted accuracy by up to 37 points, showing how much they were missing on their own. Another 31.8% comes from models failing to reason correctly over evidence they already have, since even the strongest model topped out at 69.1% accuracy on the hardest tasks when given gold evidence outright. The remaining roughly 11% of failures falls outside those two categories, per the researchers, and isn't itemized further.\n\nThat's a real caveat for anyone building \"AI scientist\" tools on today's benchmarks, which mostly grade the destination and ignore whether the model actually took the road to get there.","[\"ai\",\"benchmarks\",\"multimodal ai\",\"research\"]","2026-09-24T04:00:00.000Z","2026-09-24T07:48:47.027Z","2026-09-24T07:48:52.538Z","published",null,[24,30],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"publisher-r1","publisher",1,"The failure-mode breakdown is internally inconsistent — 57.2% attributed to incomplete evidence extraction plus the 'remaining' 31.8% attributed to reasoning failures with correct evidence only sums to 89%, not 100%.","resolved",{"id":31,"reviewer":32,"round":33,"reason":34,"status":29},"editor-r2","editor",2,"The failure-mode percentages still only sum to 89% (57.2% + 31.8%), not 100%, so add a clause accounting for the remaining ~11% of failures or otherwise reconcile the math before publishing.","ai",[35,37,38,39],"benchmarks","multimodal ai","research",[41],{"name":42,"url":43},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.11243",0,{"sections":46},[47,50,54,59,64,69,74,79,84,89,94,99,104,109],{"name":48,"slug":35,"count":49,"latest_published_at":18},"AI",4424,{"name":51,"slug":52,"count":53,"latest_published_at":18},"Security","security",724,{"name":55,"slug":56,"count":57,"latest_published_at":58},"Policy","policy",380,"2026-09-23T22:53:43.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":63},"Deals","deals",227,"2026-09-24T11:08:33.000Z",{"name":65,"slug":66,"count":67,"latest_published_at":68},"Hardware","hardware",174,"2026-09-24T10:10:29.000Z",{"name":70,"slug":71,"count":72,"latest_published_at":73},"Science","science",136,"2026-09-24T09:00:00.000Z",{"name":75,"slug":76,"count":77,"latest_published_at":78},"Consumer Tech","consumer-tech",116,"2026-09-24T00:51:49.000Z",{"name":80,"slug":81,"count":82,"latest_published_at":83},"Software","software",85,"2026-09-23T20:00:00.000Z",{"name":85,"slug":86,"count":87,"latest_published_at":88},"Dev Tools","dev-tools",79,"2026-09-22T22:21:13.000Z",{"name":90,"slug":91,"count":92,"latest_published_at":93},"Startups","startups",66,"2026-09-23T17:28:38.000Z",{"name":95,"slug":96,"count":97,"latest_published_at":98},"Gaming","gaming",45,"2026-09-22T15:35:06.000Z",{"name":100,"slug":101,"count":102,"latest_published_at":103},"General","general",43,"2026-09-21T23:48:56.000Z",{"name":105,"slug":106,"count":107,"latest_published_at":108},"Reviews","reviews",27,"2026-09-22T13:00:00.000Z",{"name":110,"slug":111,"count":112,"latest_published_at":113},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]