A new benchmark finds that today's leading AI models can barely pass a real test of mathematical creativity.
Researchers built ALPS (Austin-Law Proof-Synthesis) to test whether AI outputs are original and provably correct, not just plausible-sounding. Each task asks a model to construct an infinite mathematical structure satisfying a given equational law, or prove no such structure can exist, with every submission checked by automated proof software rather than a human grader. A public generator can produce unlimited new problems, so models cannot lean on memorized answers. Across a pool of 4,141 laws, a portfolio of eight automated prover configurations solved just 2.2 percent, a twentyfold increase in compute budget added only 0.6 percent more, and the strongest reasoning model tested solved 14 percent of the proof-only half and none of the construction half.
That gap is the real story: more compute barely helped, which points to a missing method for building tailored structures rather than a raw-power problem. It is also a sharper yardstick than most AI creativity claims, which typically rely on human judges or tasks a model could have already seen during training.
Next time a lab claims its model "discovered" a new proof, this is the kind of test worth asking whether it could pass: 97.2 percent of ALPS problems remain unsolved at every configuration and budget tried so far.