[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-why-the-standard-ai-training-reward-backfires-on-small-models":10,"sections":34},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":29,"feedback":33,"feedback_at":22,"cost_usd":33,"total_tokens":33},7777,"why-the-standard-ai-training-reward-backfires-on-small-models","Why the Standard AI Training Reward Backfires on Small Models","A new study shows the standard AI training reward, borrowed from math and coding, is the worst choice for teaching small models to search and answer questions.","A new study finds that the reward function borrowed from math and coding is actually the worst way to train small AI models to search the web.\n\nResearchers trained Qwen3.5-0.8B, a model with well under a billion parameters, to search Wikipedia and answer open-domain questions using reinforcement learning with verifiable rewards (RLVR), the technique behind recent gains in math and code models. They tested three different reward designs, including the exact-match-only scoring used in the earlier Search-R1 method, training each version with Group Relative Policy Optimization across three separate runs. The best reward setup pushed average exact-match accuracy to 0.352 on a seven-benchmark question-answering suite, up from an untrained baseline of 0.092 - a 3.8-fold gain achieved with no larger teacher model involved. The exact-match-only reward finished last in every run, even on the exact-match metric it was directly optimizing for.\n\nThat result cuts against a common assumption in AI research: that a reward recipe proven on large models will simply work at smaller scale. Here, the reward shape considered a safe default for math and coding backfired specifically because the model was small, not despite it. For anyone trying to build cheap, fast search agents without a giant model to distill from, that is a concrete warning about which shortcuts do not transfer.\n\nIt is a small, narrow study - one model family, one training dataset - but it is a useful check on the assumption that scaling AI training tricks down is as simple as scaling the model down.","[\"ai\",\"reinforcement-learning\",\"small-models\",\"search-agents\"]","2026-09-25T04:00:00.000Z","2026-09-25T21:49:11.252Z","2026-09-25T21:49:17.617Z","published",null,[],"ai",[24,26,27,28],"reinforcement-learning","small-models","search-agents",[30],{"name":31,"url":32},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.28765",0,{"sections":35},[36,40,45,50,55,60,64,69,74,79,84,89,93,98],{"name":37,"slug":24,"count":38,"latest_published_at":39},"AI",4635,"2026-09-26T23:42:06.000Z",{"name":41,"slug":42,"count":43,"latest_published_at":44},"Security","security",751,"2026-09-26T12:00:00.000Z",{"name":46,"slug":47,"count":48,"latest_published_at":49},"Policy","policy",396,"2026-09-26T18:45:15.000Z",{"name":51,"slug":52,"count":53,"latest_published_at":54},"Deals","deals",260,"2026-09-26T09:00:00.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Hardware","hardware",186,"2026-09-26T17:26:54.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":54},"Science","science",144,{"name":65,"slug":66,"count":67,"latest_published_at":68},"Consumer Tech","consumer-tech",135,"2026-09-26T14:30:00.000Z",{"name":70,"slug":71,"count":72,"latest_published_at":73},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":75,"slug":76,"count":77,"latest_published_at":78},"Dev Tools","dev-tools",84,"2026-09-26T04:20:58.000Z",{"name":80,"slug":81,"count":82,"latest_published_at":83},"Startups","startups",76,"2026-09-25T18:33:59.000Z",{"name":85,"slug":86,"count":87,"latest_published_at":88},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":90,"slug":91,"count":87,"latest_published_at":92},"General","general","2026-09-26T17:02:42.000Z",{"name":94,"slug":95,"count":96,"latest_published_at":97},"Reviews","reviews",30,"2026-09-24T20:07:31.000Z",{"name":99,"slug":100,"count":101,"latest_published_at":102},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]