A new paper argues that measuring whether AI can do real mathematics requires splitting "creativity" into distinct, non-interchangeable skills, not one aggregate score.
The paper, posted to arXiv on August 18, 2026, proposes that mathematical creativity breaks into several mechanistically different modes: reflexive introspection on mathematical practice, analogical borrowing from the sciences, problem-driven construction, and bridging distant domains, plus a cross-cutting split in how conjectures form - meaning pursued because a pattern was observed versus meaning pursued because it is strategically wanted. The author argues these modes are likely non-substitutable, meaning competence in one does not transfer to the others. Drawing on historical case studies and an architecture-level look at transformer systems, the paper contends today's models concentrate their competence in modes built on recombination and search over existing building blocks - and if that holds, the other modes may be out of reach in principle, not just achievable more slowly.
That framing matters because proof-generation is getting cheaper as AI improves at it, which the paper says is already being flagged by leading voices in the field. As that happens, mathematical value shifts toward exactly the modes current systems reportedly cannot perform, which means benchmarks that produce one aggregate "math ability" score may be measuring the wrong thing.
It is a useful check on the instinct to treat a high score on competition-style problems as proof a model can invent new mathematics - a distinction leaderboards rarely bother to make.