A new benchmark asked AI models to make headway on problems nobody has ever solved, and the best model only got partway there 14 percent of the time.
Researchers built OpenProblemBench from 82 unresolved problems pulled from the math and theoretical physics literature. Each problem includes the background, the assumptions researchers already accept, and however far anyone has gotten, so models aren't starting from zero. Since there's no answer key for a problem nobody has solved, four separate AI models grade each submission on correctness, completeness, and how much real progress it makes. Across seven model configurations tested, GPT-6-Astra posted the highest average judged solve rate at 14.0 percent, full-size open-source models scored 5.5 to 6.7 percent, and smaller Flash-tier models managed only 2.4 to 3.7 percent.
That gap matters because it separates AI's long-running party trick, reciting known facts, from something closer to research: generating a correct answer where none exists in any training set. The size difference between the top score and the open-model scores also suggests raw scale is still buying real advantage on frontier reasoning problems, not just better test-taking.
Still, judged progress here is decided by other AI models grading each other's homework, not working mathematicians, so a 14 percent solve rate reads less like an AI closing in on an unsolved conjecture and more like a new leaderboard for a notably strange video game.