AI/ ai · mathematics · benchmarks · open-source

Language Models Just Solved 147 Open Math Conjectures

A new open benchmark shows large language models can crack genuine unsolved OEIS conjectures for as little as $50 a try, though extra tools barely help.

A new benchmark just showed language models can resolve genuine unsolved math conjectures, not just textbook problems.

Researchers built OEIS Open, a benchmark of 492 open conjectures drawn from the Online Encyclopedia of Integer Sequences and formalized in the Lean proof language by Tsoukalas and colleagues. Unlike earlier work that tested one bespoke, custom-built agent, the new evaluation code is open-source, runs any generic language model, and is designed to resist models gaming the test. Given a minimal toolset and a budget of $50 per attempt, models resolved 147 of the 492 conjectures, a 30% hit rate. A cheaper 100-conjecture subset called OEIS Open Lite let the best current model reach 44% when given a larger $200 budget per attempt.

This is autonomous progress on real open math problems, not a benchmark built to flatter language models. But the researchers are upfront that most of these conjectures are obscure integer-sequence puzzles that had drawn little to no prior attention, so clearing them is nowhere near cracking a famous unsolved problem.

Notably, giving models access to 476,000 arXiv papers or fancier agent loops did not improve scores on OEIS Open Lite, suggesting the bottleneck is reasoning, not access to the literature. That is a useful, if humbling, data point for anyone pitching an 'AI mathematician' just yet.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →