[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-ppo-critics-go-flat-and-it-hurts-llm-training":10,"sections":35},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":30,"feedback":34,"feedback_at":22,"cost_usd":34,"total_tokens":34},6635,"ppo-critics-go-flat-and-it-hurts-llm-training","PPO Critics Go Flat, and It Hurts LLM Training","A new arXiv paper finds PPO value critics smooth over sharp reward changes, and a three-state sparse fix outperforms the standard approach on Qwen3.","Reinforcement learning has a quiet math problem, and researchers think they found it.\n\nA paper posted to arXiv describes \"Value Flattening,\" a failure mode in Proximal Policy Optimization, the algorithm most large language models use for RL fine-tuning. The critic component of PPO is supposed to estimate how good a given state is, so the policy knows which moves to reinforce. The researchers found that when they measured true state values using Monte Carlo rollouts, those values swung sharply between intermediate states. The critic's own predictions, by contrast, stayed comparatively flat. They confirmed the effect in a controlled FrozenLake test and found it worsens as the state space grows, then traced the cause to an implicit variance penalty in the critic's loss function plus redundant updates from similar, temporally correlated states.\n\nTheir fix, called SP3O, is almost aggressively simple: instead of applying the value loss to every state in a response, apply it to just a few well-separated ones. Tested on Qwen3-Base with only three supervised states per response, it mitigated Value Flattening and improved policy performance consistently across model sizes and evaluation suites.\n\nThis matters because PPO is the backbone of RLHF-style training across the industry, and a critic that quietly smooths over the signal it's supposed to sharpen is the kind of bug that costs compute without ever throwing an error. If sparse supervision generalizes beyond Qwen3, it's a cheap lever labs training reasoning and agentic models could pull immediately, since it changes which states get a loss, not the underlying algorithm.\n\nRL fine-tuning keeps producing these unglamorous but consequential fixes - GRPO dropped the critic model entirely; this one just tells it to look at less.","[\"reinforcement-learning\",\"llm-training\",\"ppo\",\"machine-learning\"]","2026-09-17T04:00:00.000Z","2026-09-18T03:38:36.387Z","2026-09-18T03:38:48.407Z","published",null,[],"ai",[26,27,28,29],"reinforcement-learning","llm-training","ppo","machine-learning",[31],{"name":32,"url":33},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.18708",0,{"sections":36},[37,41,45,50,55,59,63,68,73,77,82,87,92,97],{"name":38,"slug":24,"count":39,"latest_published_at":40},"AI",3852,"2026-09-17T08:27:09.000Z",{"name":42,"slug":43,"count":44,"latest_published_at":18},"Security","security",648,{"name":46,"slug":47,"count":48,"latest_published_at":49},"Policy","policy",338,"2026-09-11T04:00:00.000Z",{"name":51,"slug":52,"count":53,"latest_published_at":54},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":18},"Hardware","hardware",154,{"name":60,"slug":61,"count":62,"latest_published_at":18},"Science","science",114,{"name":64,"slug":65,"count":66,"latest_published_at":67},"Consumer Tech","consumer-tech",99,"2026-09-09T17:27:33.000Z",{"name":69,"slug":70,"count":71,"latest_published_at":72},"Software","software",75,"2026-09-10T20:41:21.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":18},"Dev Tools","dev-tools",73,{"name":78,"slug":79,"count":80,"latest_published_at":81},"Startups","startups",55,"2026-09-09T23:14:29.000Z",{"name":83,"slug":84,"count":85,"latest_published_at":86},"Gaming","gaming",43,"2026-09-10T12:18:06.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"General","general",41,"2026-09-08T01:57:23.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]