Qwen3 can solve competition math problems in any language. It just won't show its work outside English.
Researchers audited thirteen endpoints across the Qwen3 model family on competition-mathematics problems in eleven languages, checking whether each attempt actually finished with a correct, visible chain of reasoning in the language the user asked in. In English, 92.9% of problems got a complete, correct solution. In the ten other languages, that number dropped to 15.4-17.9%, even across sixteen attempts per problem. Multilingual supervised fine-tuning brought reasoning back into the target language, but it also lowered accuracy and made non-English reasoning traces more prone to looping forever without finishing. Reinforcement learning fixed that non-termination problem across the board, but only the version whose reward function explicitly scored language match kept answers in the requested language - rewarding correctness alone just pulled the model back into English.
This isn't a benchmark curiosity. A model that can solve a problem but only delivers the answer in English isn't giving non-English speakers a weaker experience - it's giving them no access to the same capability. That distinction matters for anyone deploying open reasoning models outside English-speaking markets, and it points to reward design, not training-data volume, as the lever that actually closes the gap.
The fix here isn't "add more languages to the pretraining mix" - it's "tell the reward function to care," a cheaper and more surgical move than the field's usual instinct to just scale up multilingual data.