[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-small-model-learns-to-spot-fake-math-theorems-via-rl":10,"sections":34},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":29,"feedback":33,"feedback_at":22,"cost_usd":33,"total_tokens":33},9988,"small-model-learns-to-spot-fake-math-theorems-via-rl","Small Model Learns to Spot Fake Math Theorems via RL","A 4B parameter model trained with reinforcement learning learns to disprove false math theorems, reversing a collapse caused by imitation-only fine-tuning.","A new training method teaches small AI models the thing they're worst at: proving a theorem is false.\n\nResearchers built SymCE, a set of 4,707 deliberately false undergraduate-algebra and real-analysis conjectures, each checked by its own Python verifier that doubles as a training reward. They used it to train Qwen3-4B, a small open model, first with standard supervised fine-tuning on counterexamples, then with reinforcement learning (GRPO) using only pass-fail signals from the verifier. The supervised-only version got worse at recognizing true theorems, dropping from 27% accuracy to zero. The reinforcement-learned version fixed that collapse and pushed accuracy to 66%, a result that held across four random seeds and on a second model, Gemma-3-4B.\n\nThe finding undercuts the common assumption that imitation learning is a safe default before reinforcement learning. Here, fine-tuning on a narrow skill - generating counterexamples - didn't just fail to help with a related skill, verifying true theorems. It destroyed it. Sparse, outcome-only rewards avoided that trap in a way denser partial-credit rewards did not, per a 33-point gap on a held-out probe. The resulting 4B model beat every 7B open-weights math specialist tested and stayed close to six commercial frontier APIs.\n\nIt's a narrow, specific result, but it's a clean data point in the broader argument that for reasoning tasks, learning from your own mistakes beats memorizing someone else's answers.","[\"ai\",\"reinforcement-learning\",\"llm-training\",\"math-reasoning\"]","2026-10-05T04:00:00.000Z","2026-10-05T16:50:41.167Z","2026-10-05T16:50:45.958Z","published",null,[],"ai",[24,26,27,28],"reinforcement-learning","llm-training","math-reasoning",[30],{"name":31,"url":32},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.02444",0,{"sections":35},[36,39,43,48,53,58,62,67,71,75,80,85,90,95],{"name":37,"slug":24,"count":38,"latest_published_at":18},"AI",6233,{"name":40,"slug":41,"count":42,"latest_published_at":18},"Security","security",868,{"name":44,"slug":45,"count":46,"latest_published_at":47},"Policy","policy",444,"2026-10-03T15:02:01.000Z",{"name":49,"slug":50,"count":51,"latest_published_at":52},"Deals","deals",323,"2026-10-04T13:00:00.000Z",{"name":54,"slug":55,"count":56,"latest_published_at":57},"Hardware","hardware",204,"2026-10-03T14:50:50.000Z",{"name":59,"slug":60,"count":61,"latest_published_at":18},"Science","science",177,{"name":63,"slug":64,"count":65,"latest_published_at":66},"Consumer Tech","consumer-tech",158,"2026-10-03T03:21:12.000Z",{"name":68,"slug":69,"count":70,"latest_published_at":18},"Dev Tools","dev-tools",97,{"name":72,"slug":73,"count":70,"latest_published_at":74},"Software","software","2026-10-04T10:00:00.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Startups","startups",92,"2026-10-04T14:36:25.000Z",{"name":81,"slug":82,"count":83,"latest_published_at":84},"Gaming","gaming",53,"2026-10-02T02:50:39.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":89},"General","general",51,"2026-10-05T02:35:01.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"Reviews","reviews",32,"2026-10-02T18:00:00.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"How-To","how-to",7,"2026-10-01T09:00:00.000Z"]