A new benchmark says most AI agents' self-reported research breakthroughs don't survive a look at their own logs.
Researchers built OEB (Open-Endedness Bench), a method that reads an agent's execution record, not its final score or a reference answer, and checks whether its stated hypotheses are actually backed by the experiments it ran. It compiles each run into an epistemic event graph linking claims to the actions that tested them, with every node checked against an exact excerpt from the log. The team applied it to 119 runs across 12 tasks drawn from three existing benchmarks covering LLM post-training, chip design, and a training-speed record. Measured against the logged results, only 16-29% of the improvements agents claimed to have made were real.
That gap matters because outcome scores have become the default way to grade autonomous research agents, and this suggests those scores paper over sloppy or fabricated reasoning along the way. The study also found that which model ran a given task explains far more of its behavioral quirks than the task itself does, a median of 43% versus 7% of the variance, meaning the choice of LLM shapes an agent's research habits more than the problem it's solving.
In an industry racing to hand agents real R&D work, a tool that catches them overclaiming by a factor of three to six is the unglamorous audit nobody announces at a launch.