[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-test-reveals-when-ai-agents-talk-each-other-into-wrong-answers":10,"sections":34},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":29,"feedback":33,"feedback_at":22,"cost_usd":33,"total_tokens":33},9501,"new-test-reveals-when-ai-agents-talk-each-other-into-wrong-answers","New Test Reveals When AI Agents Talk Each Other Into Wrong Answers","A new audit framework shows that letting AI agents share full reasoning, not just answers, boosts both correct fixes and newly introduced errors.","A new auditing framework shows that letting AI agents swap full reasoning, not just final answers, makes them better at fixing each other's mistakes, and just as prone to talking each other into new ones.\n\nThe paper, posted to arXiv on October 2, introduces Independent-Communicate-Revise (ICR), a framework for evaluating multi-agent AI communication separately from raw reasoning power. ICR fixes each agent's initial reasoning trajectory, then measures whether a second agent's message causes a correction or a preservation of that answer, conditioned on whether each agent started out right or wrong in the first place. It also runs a no-message control, so any extra revision can be credited to the communication itself rather than just additional inference. The audit covered four reasoning benchmarks, including MedQA and GPQA-D, and tested both text messages and latent, non-text communication channels.\n\nThe headline finding: full reasoning messages correct more wrong answers than bare answer-sharing does, but they also overturn more correct ones, on every benchmark tested. That matters because most multi-agent AI benchmarks report only final accuracy, which can hide this tradeoff entirely, since a system might post an identical score while quietly trading away correct answers for incorrect ones behind the scenes. A structured verification policy shifted agents toward preserving more answers and correcting fewer, though its effect on which answers got flagged varied by task and channel.\n\nIn other words, richer agent chatter is not an unambiguous upgrade. It is a volume knob that turns up both the right answers and the wrong ones, which should give anyone benchmarking agent collaboration pause before celebrating a higher score.","[\"ai\",\"multi-agent-systems\",\"llm-research\",\"ai-safety\"]","2026-10-02T04:00:00.000Z","2026-10-02T22:10:12.018Z","2026-10-02T22:10:18.734Z","published",null,[],"ai",[24,26,27,28],"multi-agent-systems","llm-research","ai-safety",[30],{"name":31,"url":32},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.01042",0,{"sections":35},[36,39,43,47,52,57,61,66,71,76,81,86,91,96],{"name":37,"slug":24,"count":38,"latest_published_at":18},"AI",5859,{"name":40,"slug":41,"count":42,"latest_published_at":18},"Security","security",833,{"name":44,"slug":45,"count":46,"latest_published_at":18},"Policy","policy",438,{"name":48,"slug":49,"count":50,"latest_published_at":51},"Deals","deals",317,"2026-10-01T22:00:00.000Z",{"name":53,"slug":54,"count":55,"latest_published_at":56},"Hardware","hardware",198,"2026-10-01T17:38:48.000Z",{"name":58,"slug":59,"count":60,"latest_published_at":18},"Science","science",171,{"name":62,"slug":63,"count":64,"latest_published_at":65},"Consumer Tech","consumer-tech",155,"2026-10-01T19:54:10.000Z",{"name":67,"slug":68,"count":69,"latest_published_at":70},"Dev Tools","dev-tools",96,"2026-10-01T16:57:03.000Z",{"name":72,"slug":73,"count":74,"latest_published_at":75},"Software","software",93,"2026-09-30T21:41:11.000Z",{"name":77,"slug":78,"count":79,"latest_published_at":80},"Startups","startups",90,"2026-10-01T21:55:22.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Gaming","gaming",53,"2026-10-02T02:50:39.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"How-To","how-to",7,"2026-10-01T09:00:00.000Z"]