[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-training-trick-teaches-ai-agents-to-predict-what-happens-next":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},6919,"new-training-trick-teaches-ai-agents-to-predict-what-happens-next","New Training Trick Teaches AI Agents to Predict What Happens Next","Training AI agents to predict outcomes, not just choose actions, sharpens their reinforcement-learning performance on coding and terminal tasks.","A new twist on training AI agents makes them noticeably better at learning from experience, and it doesn't require a single extra piece of data.\n\nResearchers built ActObs, a change to how agents are prepared for reinforcement learning. Standard supervised fine-tuning only trains a model to predict the actions it takes, treating everything the environment sends back as background context rather than something worth learning from. ActObs also trains the model to predict those observation tokens. Tested with the GRPO reinforcement-learning method on Terminal-Bench 2.0, agents pretrained with ActObs solved more tasks across the board on Qwen3-4B, and on Qwen3-8B they traded a little first-try reliability for a 3.4 percentage point gain at pass@16 and more distinct tasks solved overall.\n\nThe interesting part is why. The researchers trace the difference to standard fine-tuning itself: action and observation gradients quickly become orthogonal, so action-only training actually makes the model worse at predicting what happens in its environment than the untrained base model. ActObs stops that one-sided specialization, keeps more randomness in the policy, and needs smaller updates during reinforcement learning. The effect generalized too: on aider-polyglot, a code-editing benchmark neither model saw during training, the 4B model's pass@1 score jumped 4.2 percentage points.\n\nIt's a modest architectural footnote, not a new paradigm. But it's a useful reminder that what you do before reinforcement learning starts can matter as much as the RL algorithm itself.","[\"ai\",\"reinforcement-learning\",\"ai-agents\",\"research\"]","2026-09-18T04:00:00.000Z","2026-09-18T22:41:14.449Z","2026-09-18T22:41:26.370Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The aider-polyglot 4.2pp pass@1 gain is only reported in the source for the Qwen3-4B model, but the draft states it without model attribution right after two clauses that do specify model size (4B vs 8B), making it read as if the figure applies broadly — revise to explicitly tie the 4.2pp gain to the 4B model.","resolved","ai",[30,32,33,34],"reinforcement-learning","ai-agents","research",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.20715",0,{"sections":41},[42,45,49,54,59,63,67,72,76,81,86,91,96,101],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",4082,{"name":46,"slug":47,"count":48,"latest_published_at":18},"Security","security",661,{"name":50,"slug":51,"count":52,"latest_published_at":53},"Policy","policy",339,"2026-09-17T12:00:00.000Z",{"name":55,"slug":56,"count":57,"latest_published_at":58},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":18},"Hardware","hardware",155,{"name":64,"slug":65,"count":66,"latest_published_at":18},"Science","science",125,{"name":68,"slug":69,"count":70,"latest_published_at":71},"Consumer Tech","consumer-tech",99,"2026-09-09T17:27:33.000Z",{"name":73,"slug":74,"count":75,"latest_published_at":18},"Dev Tools","dev-tools",78,{"name":77,"slug":78,"count":79,"latest_published_at":80},"Software","software",75,"2026-09-10T20:41:21.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Startups","startups",55,"2026-09-09T23:14:29.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"Gaming","gaming",43,"2026-09-10T12:18:06.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"General","general",41,"2026-09-08T01:57:23.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":102,"slug":103,"count":104,"latest_published_at":105},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]