AI agents are good at predicting outcomes and bad at explaining why those outcomes happen, according to a new benchmark.
Researchers built EurekaBench, a cross-domain test covering 26 long-horizon tasks in neuroscience, computer science, chemistry, astrophysics, geophysics, and plasma physics. Each task hands an agent raw observational data and asks it to run its own experiments, then infer the mechanism behind what it sees. The benchmark checks the resulting mechanisms against 306 expert-verified scientific insights, scoring agents on three axes: whether they respect known scientific constraints, how accurately their mechanism predicts new data, and whether it actually produces usable scientific insight.
The gap between those axes is the real finding. Agents frequently beat human scientists on raw predictive accuracy - essentially fitting a model to the data - but fall well short of humans when it comes to extracting insight, the kind of conceptual leap that let Newton connect a falling apple to an orbiting moon. That distinction matters because most "AI for science" pitches promise discovery, not just better curve-fitting.
A model that predicts well without explaining anything is a regression with better PR.