AI/ ai · llm-benchmarks · multilingual-nlp · math-reasoning

New Test Shows AI Math Reasoning Falters Across Languages

A new benchmark reveals many AI models solve the same math problem inconsistently once names, digits, and language change.

AI models that ace grade-school math often stumble when the same problem shows up in a different language or with swapped numbers.

Researchers behind the paper "MGSM-Pro: A Simple Strategy for Robust Multilingual Mathematical Reasoning Evaluation" (arXiv:2601.21225) built a tougher version of MGSM, an existing multilingual math benchmark. They generated five variants of each question, swapping names, digits, and irrelevant context, then tested models across nine languages. Low-resource languages saw sharp accuracy drops on digit swaps, even when the same models stayed steady in high-resource languages. The pattern held for proprietary systems too: Gemini 2.5 Flash and GPT-4.1 both lost ground on digit changes, while Gemini 3.0 Pro and open models GPT-OSS 120B and DeepSeek v3 proved more consistent.

That distinction matters because most benchmark leaderboards report a single score per language, not five. If changing a number from 12 to 47 flips a model's answer, it wasn't reasoning about the problem in the first place; it was pattern-matching to digits it had memorized. The researchers' fix is blunt but sensible: test each question with at least five digit variations before trusting the score.

It's a multilingual echo of an English-only finding from GSM-Symbolic, which found similar instability when questions were reworded. Robustness, it turns out, doesn't travel well between languages any more than it travels between phrasings.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →