[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-a-frozen-critic-can-replace-reward-labels-in-llm-training":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},8620,"a-frozen-critic-can-replace-reward-labels-in-llm-training","A Frozen Critic Can Replace Reward Labels in LLM Training","A new arXiv paper (2609.37119) shows a frozen critic can match PPO's results without reward labels, cutting compute for long reasoning tasks.","A new training method throws out reward labels entirely and still matches a labeled baseline.\n\nResearchers posted \"Unlocking the Critic: Reward-Free Policy Optimization for LLM Post-Training\" to arXiv on September 30, 2026, as arXiv:2609.37119, cross-listed in cs.AI. The paper pushes back on a recent trend: stripping the critic model out of reinforcement learning post-training to cut memory use and training instability. The authors argue a well-trained critic - one that predicts whether an unfinished answer will eventually succeed - is more valuable kept than thrown away once training ends. Their method, Reward-Free Policy Optimization (RFPO), freezes a single calibrated critic and reuses it three ways: as the reward for a finished rollout, as a baseline for advantage estimates, and as a forecaster that scores incomplete prefixes before generation is done. Binarizing the critic's score, the authors say, stops the policy from gaming the critic's bias toward longer outputs, and the resulting system matches supervised PPO - the standard reinforcement-learning baseline that depends on labeled reward data - without a single label in the training loop.\n\nFor long chain-of-thought tasks, where a model can generate thousands of tokens before anyone knows if the answer is right, waiting for a full rollout to score it is expensive. RFPO scores trajectories before they finish, so training no longer pays to wait on every generation - which the authors say cuts both compute and memory overhead. That is a direct challenge to two years of RL post-training work that has treated critics as the unstable, memory-hungry piece worth removing.\n\nIf independent replication holds up, it flips a design assumption baked into a lot of today's RLHF tooling: that a critic is dead weight once training ends, rather than a forecasting tool worth keeping.","[\"ai\",\"reinforcement-learning\",\"llm-training\",\"research\"]","2026-09-30T04:00:00.000Z","2026-09-30T15:24:05.605Z","2026-09-30T15:24:11.065Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Add explicit source attribution (arXiv ID\u002Fdate, e.g. arXiv:2609.37119, cs.AI cross-listing) since the piece never names where or when this came from, and replace the vague 'matches PPO' \u002F 'cutting compute and memory overhead' claims with the actual comparison figures or benchmark tasks from the paper rather than leaving the performance claim unquantified.","resolved","ai",[30,32,33,34],"reinforcement-learning","llm-training","research",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.37119",0,{"sections":41},[42,45,49,53,58,63,68,73,78,82,87,92,97,102],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",5146,{"name":46,"slug":47,"count":48,"latest_published_at":18},"Security","security",788,{"name":50,"slug":51,"count":52,"latest_published_at":18},"Policy","policy",417,{"name":54,"slug":55,"count":56,"latest_published_at":57},"Deals","deals",284,"2026-09-29T21:00:00.000Z",{"name":59,"slug":60,"count":61,"latest_published_at":62},"Hardware","hardware",194,"2026-09-29T13:16:04.000Z",{"name":64,"slug":65,"count":66,"latest_published_at":67},"Science","science",154,"2026-09-28T13:19:18.000Z",{"name":69,"slug":70,"count":71,"latest_published_at":72},"Consumer Tech","consumer-tech",142,"2026-09-29T18:38:03.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":77},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":79,"slug":80,"count":81,"latest_published_at":18},"Dev Tools","dev-tools",90,{"name":83,"slug":84,"count":85,"latest_published_at":86},"Startups","startups",83,"2026-09-29T21:51:36.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"General","general",49,"2026-09-28T16:44:57.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]