AI/ ai · benchmarks · multimodal ai · research

New Benchmark Catches AI Models Faking Scientific Reasoning

Sci-MMR shows leading multimodal AI often gets the right answer without ever assembling or reasoning through the actual evidence.

A new benchmark finds that top AI research assistants often land on the right scientific answer while skipping the evidence trail that's supposed to get them there.

Researchers built Sci-MMR, a benchmark of 235 multi-step reasoning tasks spanning four scientific disciplines, each requiring models to pull evidence from an average of nine figure panels. Rather than just checking whether the final answer is correct, the benchmark traces whether that answer is actually backed by a full chain of evidence - citations, figures, and specific highlighted regions. Testing eight frontier multimodal models, the researchers found answer accuracy beat complete-evidence recovery by more than 20 percentage points. In plain terms: the models guess right more often than they can show their work.

The gap splits into two main failure types, plus a smaller leftover the study doesn't break down further. About 57.2% of errors trace to models failing to extract complete evidence from scientific figures in the first place - handing them the correct evidence directly boosted accuracy by up to 37 points, showing how much they were missing on their own. Another 31.8% comes from models failing to reason correctly over evidence they already have, since even the strongest model topped out at 69.1% accuracy on the hardest tasks when given gold evidence outright. The remaining roughly 11% of failures falls outside those two categories, per the researchers, and isn't itemized further.

That's a real caveat for anyone building "AI scientist" tools on today's benchmarks, which mostly grade the destination and ignore whether the model actually took the road to get there.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →