[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-researchers-teach-ai-agents-to-nudge-rivals-toward-cooperation":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},8085,"researchers-teach-ai-agents-to-nudge-rivals-toward-cooperation","Researchers Teach AI Agents to Nudge Rivals Toward Cooperation","A new algorithm lets AI agents learn how much to weigh a rival's payoff while training, aiming to dodge the selfish dead ends common in multi-agent learning.","A new training method lets AI agents learn not just how to win, but how much to care about their opponents winning too.\n\nResearchers describe an algorithm called Preference-based Opponent Shaping, or PBOS, in a paper posted to arXiv. In multi-agent settings, an agent that only chases its own reward can get stuck in a bad local optimum, because every agent's payoff depends on what everyone else does. PBOS adds a \"preference parameter\" directly into an agent's loss function, so it factors in an opponent's losses as it updates its own strategy. That parameter isn't fixed - it's learned alongside the strategy itself, so an agent can shift toward cooperation or competition depending on the game it's actually in. Tested across a range of differentiable games, the method produced better reward distribution than approaches that don't model opponent preferences at all.\n\nThe interesting part is the generalization angle. Earlier opponent-modeling and opponent-shaping methods, per the paper, tend to lean on simple predictions of how a rival's strategy will change next, which works for the scenario they were built for and falls apart outside it. Letting the preference itself be learned, rather than hard-coded as \"assume cooperation\" or \"assume competition,\" is a more honest way to handle the messiness of real multi-agent training.\n\nStill, this is a differentiable-games paper, not a deployed system. Toy game environments are a long way from agents negotiating supply chains or splitting compute budgets, and the leap from \"learns a preference parameter in a controlled game\" to \"behaves predictably around other AI systems in the wild\" is exactly where these results usually get tested hardest.","[\"multi-agent-systems\",\"reinforcement-learning\",\"game-theory\",\"ai-research\"]","2026-09-28T04:00:00.000Z","2026-09-28T08:49:04.097Z","2026-09-28T08:49:10.018Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Fix the garbled text artifact in the body ('game环境' should read 'game environment') before republishing.","resolved","ai",[32,33,34,35],"multi-agent-systems","reinforcement-learning","game-theory","ai-research",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2412.03072",0,{"sections":42},[43,46,50,55,60,65,69,74,79,84,89,94,98,103],{"name":44,"slug":30,"count":45,"latest_published_at":18},"AI",4792,{"name":47,"slug":48,"count":49,"latest_published_at":18},"Security","security",762,{"name":51,"slug":52,"count":53,"latest_published_at":54},"Policy","policy",399,"2026-09-27T18:39:02.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Deals","deals",261,"2026-09-27T15:30:35.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Hardware","hardware",188,"2026-09-27T20:46:36.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":18},"Science","science",151,{"name":70,"slug":71,"count":72,"latest_published_at":73},"Consumer Tech","consumer-tech",135,"2026-09-26T14:30:00.000Z",{"name":75,"slug":76,"count":77,"latest_published_at":78},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":80,"slug":81,"count":82,"latest_published_at":83},"Dev Tools","dev-tools",84,"2026-09-26T04:20:58.000Z",{"name":85,"slug":86,"count":87,"latest_published_at":88},"Startups","startups",76,"2026-09-25T18:33:59.000Z",{"name":90,"slug":91,"count":92,"latest_published_at":93},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":95,"slug":96,"count":92,"latest_published_at":97},"General","general","2026-09-26T17:02:42.000Z",{"name":99,"slug":100,"count":101,"latest_published_at":102},"Reviews","reviews",30,"2026-09-24T20:07:31.000Z",{"name":104,"slug":105,"count":106,"latest_published_at":107},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]