[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-ai-researchers-cut-training-compute-with-smarter-critics":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},9127,"ai-researchers-cut-training-compute-with-smarter-critics","AI Researchers Cut Training Compute With Smarter Critics","AC2, a new RL method, lets a learned critic judge partial rollouts, cutting training compute 2.5x versus the standard GRPO baseline.","Researchers have found a way to train language models with reinforcement learning without waiting for every attempt to finish.\n\nThe method, called Actor-Critic with Action Chunking (AC2), assigns credit to short 10,000-token chunks of a model's output rather than waiting for a full rollout to reach its final reward. A learned critic scores the state at the end of each chunk, so the policy can update mid-trajectory. The researchers made this reliable with three tweaks: only trusting the critic on problems where it has proven accurate, feeding it a reference solution from a past successful attempt when one exists, and judging chunks large enough to be meaningful rather than single tokens. Testing on Qwen3-4B with the FineProofs-RL dataset and IMO-ProofBench benchmark, AC2 beat Group Relative Policy Optimization (GRPO)'s peak validation score of 18.5% while using 2.5 times fewer decoding FLOPs.\n\nThat efficiency gain matters because RL training for reasoning models is expensive largely because every rollout runs to completion before the model learns anything from it. AC2 needs 25% fewer training steps and generates fewer tokens per step, since it doesn't have to let weak attempts play out. If the approach generalizes beyond proof-writing benchmarks, it could meaningfully cut the compute bill for training reasoning-heavy models.\n\nLearned critics have a bad reputation in LLM training for being unreliable, and AC2's real trick is finding where to trust them rather than fixing that unreliability outright.","[\"reinforcement-learning\",\"llm-training\",\"ai-research\"]","2026-10-01T04:00:00.000Z","2026-10-01T21:20:56.837Z","2026-10-01T21:20:59.456Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Define the GRPO acronym on first use (e.g., 'Group Relative Policy Optimization') the same way AC2 is spelled out, since the piece currently names it as the baseline without ever expanding it.","resolved","ai",[32,33,34],"reinforcement-learning","llm-training","ai-research",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.39247",0,{"sections":41},[42,45,49,53,58,63,67,72,77,81,86,91,96,101],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",5572,{"name":46,"slug":47,"count":48,"latest_published_at":18},"Security","security",815,{"name":50,"slug":51,"count":52,"latest_published_at":18},"Policy","policy",430,{"name":54,"slug":55,"count":56,"latest_published_at":57},"Deals","deals",298,"2026-09-30T21:00:26.000Z",{"name":59,"slug":60,"count":61,"latest_published_at":62},"Hardware","hardware",196,"2026-09-30T13:00:00.000Z",{"name":64,"slug":65,"count":66,"latest_published_at":18},"Science","science",163,{"name":68,"slug":69,"count":70,"latest_published_at":71},"Consumer Tech","consumer-tech",149,"2026-09-30T22:57:11.000Z",{"name":73,"slug":74,"count":75,"latest_published_at":76},"Dev Tools","dev-tools",93,"2026-10-01T02:30:48.000Z",{"name":78,"slug":79,"count":75,"latest_published_at":80},"Software","software","2026-09-30T21:41:11.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Startups","startups",84,"2026-09-30T20:39:09.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"Gaming","gaming",51,"2026-09-30T16:24:30.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":102,"slug":103,"count":104,"latest_published_at":105},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]