[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-scaling-law-reveals-limits-of-ai-reward-models":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},9031,"new-scaling-law-reveals-limits-of-ai-reward-models","New Scaling Law Reveals Limits of AI Reward Models","A new scaling law shows reward model accuracy grows with the log of training comparisons, capped by a policy's divergence budget from reward hacking.","A new paper puts a number on how much preference data it takes to train a reward model that won't get gamed.\n\nResearchers studied how AI reward models - the proxy scorers used to align language models with human preferences - degrade under heavy optimization, a failure mode known as reward hacking. They found performance scales as the square root of the smaller of two quantities: the log of the number of training comparisons, or the policy's KL-divergence budget relative to a reference model. The team tested the law using a 70B parameter model to generate gold-standard feedback and smaller 0.6B to 4B models trained as proxies, and the fit explained 97 to 99 percent of the variance across model sizes, noise levels, and optimization methods including best-of-N sampling and policy tilting. The upshot: once optimization pressure exceeds what the training data can support, more pressure does not help and can actively hurt.\n\nThis matters because teams building aligned models constantly trade off between collecting more comparison data and optimizing harder against the reward model they already have. The law suggests data collection has rapidly diminishing returns - since performance tracks the log of comparisons, doubling your data budget barely moves the needle once you have a few thousand labeled pairs. The real constraint, this framework suggests, is the divergence budget, not the size of the dataset.\n\nThe researchers frame the whole exercise as a selection problem dressed up in alignment jargon - picking the best option from noisy, roughly Gaussian signals - a less glamorous explanation for reward hacking than most papers offer, but a more falsifiable one.","[\"ai-alignment\",\"reward-models\",\"scaling-laws\",\"llm-research\"]","2026-10-01T04:00:00.000Z","2026-10-01T16:38:20.773Z","2026-10-01T16:38:24.462Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Correct the formula description: the scaling law uses the log of the number of comparisons (log(M)), not the raw data budget, so 'smaller of the data budget or the drift budget' should read 'smaller of the log of the data budget or the drift budget' to avoid misstating how reward model performance actually scales with data.","resolved","ai",[32,33,34,35],"ai-alignment","reward-models","scaling-laws","llm-research",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.38526",0,{"sections":42},[43,46,50,55,60,65,69,74,79,83,88,93,98,103],{"name":44,"slug":30,"count":45,"latest_published_at":18},"AI",5488,{"name":47,"slug":48,"count":49,"latest_published_at":18},"Security","security",809,{"name":51,"slug":52,"count":53,"latest_published_at":54},"Policy","policy",429,"2026-10-01T02:26:17.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Deals","deals",298,"2026-09-30T21:00:26.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Hardware","hardware",196,"2026-09-30T13:00:00.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":18},"Science","science",162,{"name":70,"slug":71,"count":72,"latest_published_at":73},"Consumer Tech","consumer-tech",149,"2026-09-30T22:57:11.000Z",{"name":75,"slug":76,"count":77,"latest_published_at":78},"Dev Tools","dev-tools",93,"2026-10-01T02:30:48.000Z",{"name":80,"slug":81,"count":77,"latest_published_at":82},"Software","software","2026-09-30T21:41:11.000Z",{"name":84,"slug":85,"count":86,"latest_published_at":87},"Startups","startups",84,"2026-09-30T20:39:09.000Z",{"name":89,"slug":90,"count":91,"latest_published_at":92},"Gaming","gaming",51,"2026-09-30T16:24:30.000Z",{"name":94,"slug":95,"count":96,"latest_published_at":97},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":99,"slug":100,"count":101,"latest_published_at":102},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":104,"slug":105,"count":106,"latest_published_at":107},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]