A new training method turns board-game self-play into math homework, and it works.
Researchers built a framework called Self-Play Search Distillation (SPSD) that trains MuZero-style networks on board games through self-play search, records the network's preferred moves, plausible alternatives, and value estimates at each step, then converts those search records into chain-of-thought reasoning examples. Language models trained on this synthetic data never see a single math problem during training. Applied to Qwen3-4B-Base, the technique raised average scores across six math benchmarks from 24.1 to 36.6, while the model's win rate against a held-out board game jumped from 15 percent to 45 percent.
The interesting part isn't the game-playing improvement. It's the math transfer. SPSD suggests that the skill LLMs are missing isn't math-specific at all: it's structured evaluation of competing options under uncertainty, the same skill MuZero-style search already handles well. That reframes a data-scarcity problem as an engineering problem, since executable game environments can generate labeled reasoning data far more cheaply than human annotators can.
Board games are a controlled, deterministic sandbox, and math benchmarks reward similar step-by-step logic, so the transfer is plausible without being magic. Whether this holds up past a 4-billion-parameter base model or benchmarks that reward more than step-by-step logic is still an open question the paper does not answer.