AI models that ace grade-school math often stumble when the same problem shows up in a different language or with swapped numbers.
Researchers behind the paper "MGSM-Pro: A Simple Strategy for Robust Multilingual Mathematical Reasoning Evaluation" (arXiv:2601.21225) built a tougher version of MGSM, an existing multilingual math benchmark. They generated five variants of each question, swapping names, digits, and irrelevant context, then tested models across nine languages. Low-resource languages saw sharp accuracy drops on digit swaps, even when the same models stayed steady in high-resource languages. The pattern held for proprietary systems too: Gemini 2.5 Flash and GPT-4.1 both lost ground on digit changes, while Gemini 3.0 Pro and open models GPT-OSS 120B and DeepSeek v3 proved more consistent.
That distinction matters because most benchmark leaderboards report a single score per language, not five. If changing a number from 12 to 47 flips a model's answer, it wasn't reasoning about the problem in the first place; it was pattern-matching to digits it had memorized. The researchers' fix is blunt but sensible: test each question with at least five digit variations before trusting the score.
It's a multilingual echo of an English-only finding from GSM-Symbolic, which found similar instability when questions were reworded. Robustness, it turns out, doesn't travel well between languages any more than it travels between phrasings.