[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-benchmark-finds-materials-ai-often-recites-instead-of-reasons":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},8858,"new-benchmark-finds-materials-ai-often-recites-instead-of-reasons","New Benchmark Finds Materials AI Often Recites Instead of Reasons","A new benchmark called CARAT finds materials AI often recites data instead of reasoning, and shows its score gap mostly reflects missing baseline data.","A new benchmark called CARAT argues that materials science AI models often parrot structural data instead of reasoning through it.\n\nTo test this, researchers built eight matched versions of the same question-and-answer pairs, changing only how a crystal structure was described: as a plain chemical formula, as a periodic graph, or as GraphSpace, a version that names each structural relationship explicitly. On the hardest question families, GraphSpace scored 17.3 points higher than plain formulas. But turning that same scrutiny on their own benchmark, the team found GraphSpace's 19.3-point edge over a plain periodic graph splits into two very different numbers: just 1.96 points when the plain graph already contains everything needed to answer, versus 46.7 points when the plain graph omits that information entirely.\n\nThat split matters because it separates two things people routinely conflate: a model reasoning better, and a model simply being handed more complete data. Most of GraphSpace's headline advantage traces to the second case, not the first - meaning a flashy benchmark gain can really be a benchmark design artifact. For anyone citing eval scores to justify a model's \"reasoning,\" that's a reason for caution.\n\nThe researchers also stress-tested their own benchmark: a shortcut that ignores the structural link still answered four of seven \"hardened\" question families, and a frozen model kept repeating relations it was shown even after being redirected in 95.6% of paired tests - though targeted fine-tuning pushed accurate, evidence-aware answers to 99.8%, proof the flaw is fixable on purpose, not just measurable.","[\"ai\",\"benchmarks\",\"materials science\",\"research\"]","2026-10-01T04:00:00.000Z","2026-10-01T07:58:38.946Z","2026-10-01T07:58:43.886Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The claim that '46.7 of those points came from the baseline simply lacking information' frames 46.7 as a component of the 19.3-point overall margin, but 46.7 is larger than 19.3 — the source's 1.96 and 46.7 are the margins measured separately in the matched-information and missing-information subsets, not additive pieces of the 19.3 headline figure, so rewrite that sentence to describe them as separate conditional comparisons rather than a decomposition of 19.3.","resolved","ai",[30,32,33,34],"benchmarks","materials science","research",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.38340",0,{"sections":41},[42,45,50,55,60,65,70,75,80,84,89,94,99,104],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",5271,{"name":46,"slug":47,"count":48,"latest_published_at":49},"Security","security",801,"2026-09-30T22:18:23.000Z",{"name":51,"slug":52,"count":53,"latest_published_at":54},"Policy","policy",429,"2026-10-01T02:26:17.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Deals","deals",298,"2026-09-30T21:00:26.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Hardware","hardware",196,"2026-09-30T13:00:00.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":69},"Science","science",157,"2026-09-30T15:00:56.000Z",{"name":71,"slug":72,"count":73,"latest_published_at":74},"Consumer Tech","consumer-tech",149,"2026-09-30T22:57:11.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Dev Tools","dev-tools",93,"2026-10-01T02:30:48.000Z",{"name":81,"slug":82,"count":78,"latest_published_at":83},"Software","software","2026-09-30T21:41:11.000Z",{"name":85,"slug":86,"count":87,"latest_published_at":88},"Startups","startups",84,"2026-09-30T20:39:09.000Z",{"name":90,"slug":91,"count":92,"latest_published_at":93},"Gaming","gaming",51,"2026-09-30T16:24:30.000Z",{"name":95,"slug":96,"count":97,"latest_published_at":98},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":100,"slug":101,"count":102,"latest_published_at":103},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":105,"slug":106,"count":107,"latest_published_at":108},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]