A new training method teaches small AI models the thing they're worst at: proving a theorem is false.
Researchers built SymCE, a set of 4,707 deliberately false undergraduate-algebra and real-analysis conjectures, each checked by its own Python verifier that doubles as a training reward. They used it to train Qwen3-4B, a small open model, first with standard supervised fine-tuning on counterexamples, then with reinforcement learning (GRPO) using only pass-fail signals from the verifier. The supervised-only version got worse at recognizing true theorems, dropping from 27% accuracy to zero. The reinforcement-learned version fixed that collapse and pushed accuracy to 66%, a result that held across four random seeds and on a second model, Gemma-3-4B.
The finding undercuts the common assumption that imitation learning is a safe default before reinforcement learning. Here, fine-tuning on a narrow skill - generating counterexamples - didn't just fail to help with a related skill, verifying true theorems. It destroyed it. Sparse, outcome-only rewards avoided that trap in a way denser partial-credit rewards did not, per a 33-point gap on a held-out probe. The resulting 4B model beat every 7B open-weights math specialist tested and stayed close to six commercial frontier APIs.
It's a narrow, specific result, but it's a clean data point in the broader argument that for reasoning tasks, learning from your own mistakes beats memorizing someone else's answers.