AI systems that claim to have solved long-standing math problems are telling a noticeably different kind of story than the mathematicians who wrote about the same problems before them.
A new study compared public AI research accounts with the human literature on 11 mathematical problems that AI sources have reported as resolved, disproved, or substantially advanced. Researchers gathered 58 human papers addressing those same targets and built 31 within-problem comparisons, scoring each on six measures covering problem resolution, method explanation, uncertainty and boundary specification, follow-up question generation, generality, and cross-disciplinary integration. AI accounts placed more emphasis on declaring the problem resolved and linking ideas across fields. Human papers spent significantly more space explaining methods, specifying assumptions and limits, and naming open questions for other researchers to pursue.
That gap is more than a style choice. Math has historically advanced through the parts AI skips: documented methods, stated limits, and open questions that let one paper's loose thread become someone else's next result. If AI writeups close a stated problem without leaving that scaffolding behind, the field gets fewer footholds to build on, even when the result itself checks out.
The study found no meaningful difference in generality between AI and human accounts, and the pattern held even when researchers dropped each of the 11 problems from the analysis one at a time, so this looks like a structural habit, not an artifact of one flashy case.