AI/ ai · reinforcement-learning · llm-training · math-reasoning

Small Model Learns to Spot Fake Math Theorems via RL

A 4B parameter model trained with reinforcement learning learns to disprove false math theorems, reversing a collapse caused by imitation-only fine-tuning.

A new training method teaches small AI models the thing they're worst at: proving a theorem is false.

Researchers built SymCE, a set of 4,707 deliberately false undergraduate-algebra and real-analysis conjectures, each checked by its own Python verifier that doubles as a training reward. They used it to train Qwen3-4B, a small open model, first with standard supervised fine-tuning on counterexamples, then with reinforcement learning (GRPO) using only pass-fail signals from the verifier. The supervised-only version got worse at recognizing true theorems, dropping from 27% accuracy to zero. The reinforcement-learned version fixed that collapse and pushed accuracy to 66%, a result that held across four random seeds and on a second model, Gemma-3-4B.

The finding undercuts the common assumption that imitation learning is a safe default before reinforcement learning. Here, fine-tuning on a narrow skill - generating counterexamples - didn't just fail to help with a related skill, verifying true theorems. It destroyed it. Sparse, outcome-only rewards avoided that trap in a way denser partial-credit rewards did not, per a 33-point gap on a held-out probe. The resulting 4B model beat every 7B open-weights math specialist tested and stayed close to six commercial frontier APIs.

It's a narrow, specific result, but it's a clean data point in the broader argument that for reasoning tasks, learning from your own mistakes beats memorizing someone else's answers.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →