[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-ai-agents-learn-tool-planning-by-copying-ant-colonies":10,"sections":35},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":30,"feedback":34,"feedback_at":22,"cost_usd":34,"total_tokens":34},9809,"ai-agents-learn-tool-planning-by-copying-ant-colonies","AI Agents Learn Tool Planning by Copying Ant Colonies","A new training method called PhGPO uses ant-colony-style pheromone trails to help AI agents remember which tool sequences actually worked.","A new reinforcement learning method borrows a trick from ant colonies to help AI agents plan multi-step tool use.\n\nResearchers published PhGPO, short for Pheromone-Guided Policy Optimization, a training technique for AI agents that call external tools across many steps. Today's agents struggle with long chains of tool calls because the number of possible paths explodes combinatorially, and most training methods treat a successful run as a one-off reward, with nothing saved for the next attempt. PhGPO instead extracts reusable transition patterns from past successful trajectories - the paper calls this a pheromone trail - and uses it to steer future policy optimization toward tool transitions that have worked before. The researchers report improved results on long-horizon tool-planning tasks in their experiments.\n\nThis targets a real bottleneck: tool-using agents, the kind that chain API calls, searches, and code execution together, get expensive to train precisely because exploration scales combinatorially with path length. Borrowing from ant colony optimization, a decades-old heuristic for routing problems, is a sign that agent training is increasingly pulling ideas from classical search and optimization, not just scaling up bigger models.\n\nWhether that pheromone trail generalizes beyond the paper's own test tasks is the question the abstract does not answer.","[\"ai agents\",\"reinforcement learning\",\"tool use\",\"arxiv research\"]","2026-10-02T04:00:00.000Z","2026-10-03T11:30:16.921Z","2026-10-03T11:30:21.274Z","published",null,[],"ai",[26,27,28,29],"ai agents","reinforcement learning","tool use","arxiv research",[31],{"name":32,"url":33},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2602.13691",0,{"sections":36},[37,40,44,48,53,57,61,66,71,76,81,86,91,96],{"name":38,"slug":24,"count":39,"latest_published_at":18},"AI",6058,{"name":41,"slug":42,"count":43,"latest_published_at":18},"Security","security",848,{"name":45,"slug":46,"count":47,"latest_published_at":18},"Policy","policy",439,{"name":49,"slug":50,"count":51,"latest_published_at":52},"Deals","deals",317,"2026-10-01T22:00:00.000Z",{"name":54,"slug":55,"count":56,"latest_published_at":18},"Hardware","hardware",199,{"name":58,"slug":59,"count":60,"latest_published_at":18},"Science","science",176,{"name":62,"slug":63,"count":64,"latest_published_at":65},"Consumer Tech","consumer-tech",155,"2026-10-01T19:54:10.000Z",{"name":67,"slug":68,"count":69,"latest_published_at":70},"Dev Tools","dev-tools",96,"2026-10-01T16:57:03.000Z",{"name":72,"slug":73,"count":74,"latest_published_at":75},"Software","software",93,"2026-09-30T21:41:11.000Z",{"name":77,"slug":78,"count":79,"latest_published_at":80},"Startups","startups",90,"2026-10-01T21:55:22.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Gaming","gaming",53,"2026-10-02T02:50:39.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"How-To","how-to",7,"2026-10-01T09:00:00.000Z"]