AI/ ai-agents · benchmark · materials-science · arxiv

AI Agents Struggle to Operate Real Materials-Science Software

A benchmark posted to arXiv, arXiv:2609.37053, finds leading multimodal AI agents succeed on just 25% of GUI tasks in professional materials-science tools.

A new benchmark says today's AI agents can't be trusted alone in a materials science lab.

According to a paper posted to arXiv, arXiv:2609.37053, researchers built MatToolBench, a real-environment benchmark of 204 tasks spanning 10 professional materials-science tools, all run inside a Windows 11 virtual machine. Tasks cover three modes: manual GUI operation, scripting in the OriginPro plotting package, and code-based database queries. Domain experts broke each task into fine-grained sub-criteria so partial credit could be scored, and the paper reports its GUI-scoring pipeline reaches an average F1 of 0.98 against human judgment. Even the best model tested in the study managed only a 25% success rate on GUI tasks and 45% on code-based tasks.

That gap matters because it undercuts the assumption that agents good at general software are good at specialized software. The paper attributes the failures not to visual-grounding problems but to missing domain-specific operational knowledge, sparse pretraining coverage of niche scientific tools, weak handoff of artifacts between programs, and critical interface states that only show up visually.

It's a useful gut check for anyone picturing AI agents running unattended lab pipelines any time soon: an agent that aces a browser-automation benchmark can still fumble a plotting tool a materials scientist uses every day.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →