Ask an AI model to verify a materials-science claim, and it will often just make one up instead.
Researchers built a benchmark called T-MOF, covering four task types: checking whether a structure matches known identifiers, verifying synthesis conditions, judging whether cited evidence is actually sufficient, and running machine-learning interatomic potential, or MLIP, computations. They tested models in closed-book, retrieval-enabled, and oracle-evidence modes to isolate whether failures came from missing knowledge, poor evidence-gathering, or faulty reasoning. Based on what broke, they built MOF-Verify, an agentic harness that checks structural grounding, chases down literature, flags insufficient evidence, and runs computation before issuing a verdict. Across several backbone language models, MOF-Verify beat both direct inference and plain retrieval baselines.
Metal-organic frameworks are porous materials used for carbon capture, gas storage, and catalysis - exactly the chemistry labs want AI to help speed up. This benchmark shows a model can sound confident about a MOF claim while being wrong in several distinct ways: misidentifying the structure, misreading the synthesis recipe, or accepting thin evidence as proof. That undercuts the pitch that AI co-scientists have already solved verification; this paper's whole premise is that they haven't.
It's a reminder that bolting an agent onto a language model doesn't close a reasoning gap - it just adds checkpoints around it, useful only if those checkpoints catch the right failures.