[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-researchers-design-fix-for-ai-agents-that-learn-to-collude":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},6914,"researchers-design-fix-for-ai-agents-that-learn-to-collude","Researchers Design Fix for AI Agents That Learn to Collude","A new reward-shaping method called CURB stops reinforcement learning agents from silently coordinating on higher prices in repeated market interactions.","AI trading agents can quietly learn to fix prices with each other, even without ever exchanging a message. New research proposes a fix.\n\nThe problem: when reinforcement learning agents repeatedly compete, such as in pricing algorithms, they can converge on supra-competitive outcomes that look like collusion, even though no one programmed them to cooperate. Researchers formalized this behavior using Simple Penal Codes, a game-theory framework for punishment-based cooperation, and showed that any meaningful version of it leaves a measurable fingerprint: a detectable difference in how an agent acts after cooperation versus after defection. They built on that signal to create CURB (Collusion Unwinding via Reward shaping and Belief injection), which penalizes that fingerprint during training. The method is proven to eliminate the punishment threats that let collusive equilibria hold.\n\nTested in Bertrand and Cournot repeated-game simulations, two standard economic models for price and quantity competition, CURB substantially cut collusion by Q-learning agents. The researchers also extended CURB to deep Q-network agents in Bertrand competition specifically, suggesting the approach isn't limited to simple tabular learning setups.\n\nThis matters because algorithmic collusion has mostly been treated as a narrow problem for specific markets, like ad auctions or ride-hailing platforms. A general-purpose fix that works across repeated-game settings is a meaningfully different proposition: it's a tool regulators or platform operators could plausibly bolt onto pricing algorithms before deployment, rather than patching each new case after the fact.\n\nWorth noting: this is simulation work, not evidence pulled from real trading or pricing systems, and deep-network results so far cover only one of the two tested market structures.","[\"reinforcement learning\",\"algorithmic collusion\",\"game theory\",\"ai safety\"]","2026-09-18T04:00:00.000Z","2026-09-18T22:23:47.940Z","2026-09-18T22:23:59.841Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The paragraph implies the deep Q-network results held across both Bertrand and Cournot settings, but the source only reports the deep Q-network extension for Bertrand competition — rewrite that sentence to scope the deep-Q-network claim to Bertrand only.","resolved","ai",[32,33,34,35],"reinforcement learning","algorithmic collusion","game theory","ai safety",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.20548",0,{"sections":42},[43,46,50,55,60,64,68,73,77,82,87,92,97,102],{"name":44,"slug":30,"count":45,"latest_published_at":18},"AI",4082,{"name":47,"slug":48,"count":49,"latest_published_at":18},"Security","security",661,{"name":51,"slug":52,"count":53,"latest_published_at":54},"Policy","policy",339,"2026-09-17T12:00:00.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":18},"Hardware","hardware",155,{"name":65,"slug":66,"count":67,"latest_published_at":18},"Science","science",125,{"name":69,"slug":70,"count":71,"latest_published_at":72},"Consumer Tech","consumer-tech",99,"2026-09-09T17:27:33.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":18},"Dev Tools","dev-tools",78,{"name":78,"slug":79,"count":80,"latest_published_at":81},"Software","software",75,"2026-09-10T20:41:21.000Z",{"name":83,"slug":84,"count":85,"latest_published_at":86},"Startups","startups",55,"2026-09-09T23:14:29.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"Gaming","gaming",43,"2026-09-10T12:18:06.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"General","general",41,"2026-09-08T01:57:23.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]