[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-a-new-controller-for-ai-post-training-beats-its-best-rival":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},9907,"a-new-controller-for-ai-post-training-beats-its-best-rival","A new controller for AI post-training beats its best rival","A new budget-aware RL controller, FSPO, edges past the strongest prior baseline on accuracy while cutting miscalibration and blocking infeasible moves.","Researchers built a smarter traffic controller for AI training, and it squeezes a bit more accuracy out of the same compute budget.\n\nFSPO is a feedback-state controller for reinforcement-learning post-training of large language models, the kind of training that adjusts knobs like rollout temperature, group size, and verifier allocation while a model trains under a fixed compute budget. It replaces ad hoc tuning with a risk model trained to match the exact controller making future decisions, a calibration method called DCTC that recalibrates risk estimates as the controller's own choices shift the data it sees, and a feasibility checker called PRCC that blocks any action that would leave the remaining budget unable to finish the run. Tested under a matched compute budget against PB2, the strongest adaptive baseline in the paper, FSPO reached 66.11% held-out accuracy and 59.43% out-of-distribution accuracy, versus 64.47% and 57.03% for PB2. The same components also cut calibration error (ECE) from 0.108 to 0.053 and, on an 18-action catalog, eliminated false-feasible admissions that previously happened 19.7% of the time.\n\nRL post-training runs are expensive, and most resource controllers either tune knobs with simple heuristics or risk promising more compute than they can actually deliver mid-run. FSPO's real contribution is structural: it targets specific failure modes, like a risk model that doesn't match the controller it's paired with, or a budget plan that quietly becomes infeasible, rather than just chasing a bigger accuracy number.\n\nThe accuracy edge over PB2 is real but modest, 1.64 points held-out and 2.40 points out-of-distribution, so the more convincing result here is operational: trajectory failures dropped from 18.1% to 8.3% once all three components were switched on.","[\"reinforcement-learning\",\"llm-post-training\",\"ai-research\",\"compute-budget\"]","2026-10-05T04:00:00.000Z","2026-10-05T12:50:22.759Z","2026-10-05T12:50:27.276Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Fix the closing line's math: the article's own numbers put FSPO's gain over PB2 at 1.64 and 2.40 percentage points (66.11-64.47 and 59.43-57.03), not 'two to three' as stated — the +2.42\u002F+3.19 figures in the source are versus the contextual bandit baseline, not PB2, and shouldn't be conflated with the PB2 comparison.","resolved","ai",[32,33,34,35],"reinforcement-learning","llm-post-training","ai-research","compute-budget",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.02828",0,{"sections":42},[43,46,50,55,60,65,69,74,78,82,87,92,97,102],{"name":44,"slug":30,"count":45,"latest_published_at":18},"AI",6166,{"name":47,"slug":48,"count":49,"latest_published_at":18},"Security","security",859,{"name":51,"slug":52,"count":53,"latest_published_at":54},"Policy","policy",444,"2026-10-03T15:02:01.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Deals","deals",323,"2026-10-04T13:00:00.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Hardware","hardware",204,"2026-10-03T14:50:50.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":18},"Science","science",177,{"name":70,"slug":71,"count":72,"latest_published_at":73},"Consumer Tech","consumer-tech",158,"2026-10-03T03:21:12.000Z",{"name":75,"slug":76,"count":77,"latest_published_at":18},"Dev Tools","dev-tools",97,{"name":79,"slug":80,"count":77,"latest_published_at":81},"Software","software","2026-10-04T10:00:00.000Z",{"name":83,"slug":84,"count":85,"latest_published_at":86},"Startups","startups",92,"2026-10-04T14:36:25.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"Gaming","gaming",53,"2026-10-02T02:50:39.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"General","general",51,"2026-10-05T02:35:01.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"Reviews","reviews",32,"2026-10-02T18:00:00.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"How-To","how-to",7,"2026-10-01T09:00:00.000Z"]