AI/ ai-agents · ai-research · benchmarks · autonomous-agents

AI Research Agents Still Mostly Reuse Old Tricks

A study testing seven frontier AI models on 36 long research tasks finds they mostly recombine known techniques instead of inventing new ones.

A new evaluation of AI agents doing real research and development work finds they're better engineers than scientists.

Researchers tested seven frontier models on 36 long-horizon R&D tasks using a framework that scores behavior in three parts - solution framing, execution, and feedback control - instead of just a final result. They also ran controlled comparisons to see whether agents actually get better at using their own past experience, both within a single task and across different ones. The agents could formulate and carry out workable solutions, but their performance swung widely from run to run. Even their strongest results were mostly adaptations or combinations of established techniques; real methodological novelty was rare.

That matters because it complicates the pitch that AI agents are inching toward autonomous scientific discovery. The paper's more useful finding is subtler: two agents can land on the same final score for completely different reasons, hitting different bottlenecks along the way, and "remembering" past experience can steer later decisions wrong just as easily as it helps. The researchers also found that harness design - how the agent's tools and workflow are structured - measurably affects how stable its performance is.

So treat the current crop of research agents as capable optimizers, not junior scientists. They're good at combining what already works. Whether they can reliably originate what doesn't yet exist is still an open question, and this paper is one of the more careful attempts to actually measure that gap instead of assuming it away.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →