AI/ ai · ai-agents · materials-science · benchmarks

Benchmark Tests Whether AI Agents Can Do Real Materials Science

A 94-task benchmark grades AI agents on real computational materials science steps, and even top models stumble on long, loosely guided workflows.

Researchers have built a benchmark that tests AI agents on real materials-science research tasks without re-running the expensive simulations themselves.

The new benchmark, called CompMat-Bench, pulls 94 tasks from published computational materials studies. Each task asks an agent to complete one concrete research step, like preparing simulation inputs or analyzing the results, rather than running the simulation itself. The team pre-computed the correct inputs and outputs ahead of time, so grading relies on fixed rules instead of a costly re-run or an LLM acting as judge. Agents built on three different language models were tested across single tasks and multi-step workflows, with full or reduced methodological guidance.

On single tasks with full guidance, the agents passed 66.0 to 90.4 percent of the time, showing current models can handle isolated steps of real science reasonably well. Stack those steps into a longer workflow and strip away the hand-holding, though, and most agents falter, with one exception: the strongest model only breaks down when both problems hit at once. The more telling result is in the failure analysis, agents aren't tripping over software, they're making scientific errors.

That distinction matters more than another leaderboard score. It is the difference between an agent that can follow a recipe and one that understands the chemistry behind it.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →