AI/ ai · benchmarks · materials science · research

New Benchmark Finds Materials AI Often Recites Instead of Reasons

A new benchmark called CARAT finds materials AI often recites data instead of reasoning, and shows its score gap mostly reflects missing baseline data.

A new benchmark called CARAT argues that materials science AI models often parrot structural data instead of reasoning through it.

To test this, researchers built eight matched versions of the same question-and-answer pairs, changing only how a crystal structure was described: as a plain chemical formula, as a periodic graph, or as GraphSpace, a version that names each structural relationship explicitly. On the hardest question families, GraphSpace scored 17.3 points higher than plain formulas. But turning that same scrutiny on their own benchmark, the team found GraphSpace's 19.3-point edge over a plain periodic graph splits into two very different numbers: just 1.96 points when the plain graph already contains everything needed to answer, versus 46.7 points when the plain graph omits that information entirely.

That split matters because it separates two things people routinely conflate: a model reasoning better, and a model simply being handed more complete data. Most of GraphSpace's headline advantage traces to the second case, not the first - meaning a flashy benchmark gain can really be a benchmark design artifact. For anyone citing eval scores to justify a model's "reasoning," that's a reason for caution.

The researchers also stress-tested their own benchmark: a shortcut that ignores the structural link still answered four of seven "hardened" question families, and a frozen model kept repeating relations it was shown even after being redirected in 95.6% of paired tests - though targeted fine-tuning pushed accurate, evidence-aware answers to 99.8%, proof the flaw is fixable on purpose, not just measurable.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →