A new benchmark says today's AI agents are nowhere close to doing science on their own.
Researchers built ASI-Bench, the first benchmark to test both innovative exploration and autonomous scientific execution in AI systems. It took more than 40 experts and over 31,000 human hours to build 60 project-level research tasks spanning 11 scientific domains. Unlike most benchmarks, it progressively strips away methodological guidance within the same research project to see how far a model can get on its own. Every task passes through expert review, AI-assisted auditing, sandbox execution, and scorer validation before a score counts. Across 18 leading agent-model combinations, average scores fell from 50.91 with full guidance to 29.10 once only the method was specified, and down to 26.62 when the agent had to pick its own method.
That drop matters because it separates two very different skills: applying a method someone hands you, and inventing one yourself. Current systems are clearly good at the first and still weak at the second, which is precisely the gap between a capable research assistant and the open-ended, self-directed science that superintelligence claims require.
Every lab talks up AI that will make discoveries on its own; this benchmark suggests that ambition is still mostly a roadmap slide.