AI/ ai-agents · benchmarks · scientific-computing · computer-use-agents

Benchmark Finds AI Agents Fumble Scientific Software Tasks

A new 146-task benchmark shows even top AI models struggle to complete real scientific workflows like molecular drawing and statistical analysis.

A new benchmark says today's AI agents still can't reliably run a lab's software stack.

Researchers built OSWorld-Science, a benchmark of 146 tasks testing whether agents built on vision-language models can actually operate scientific software, not just talk about it. The tasks span molecular drawing and retrosynthesis, pathology image analysis, statistical computing, and physical simulation, and were developed with input from domain experts rather than pulled off the shelf. Instead of grading on vibes, the evaluators inspect the actual artifacts an agent produces, such as a molecular structure, a segmentation mask, a plot, or a numerical result, and hand out partial credit for near-misses. The team ran 12 vision-language models through a purpose-built harness that logs every step of the agent's interaction loop.

The finding that matters: even well-equipped, state-of-the-art models still struggle with tasks a competent lab tech would treat as routine. That's worth noting as software vendors rush to bolt agent features onto scientific tools, since it suggests the gap between a slick chat demo and dependably running an experiment is still wide. The researchers also tie performance to factors like reasoning effort and context length, which matters for anyone deciding how much compute these agents are worth.

File it next to every other computer-use benchmark that promised agents could handle any app: the demo is always smoother than the real interface.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →