AI/ ai · multi-agent systems · ai governance · trust

AI Agents Learn to Trust Teammates - and Forget to Verify

A cooperative survival game reveals frontier AI models cut teammate verification by up to 85% - and rebuild trust far more slowly than they form it.

New research gives AI governance a rare measurable thing: how much one AI agent trusts another, and what happens when that trust breaks.

A group of researchers built a cooperative survival game in which an AI agent can verify a teammate's answer - at a resource cost - or simply accept it. Getting it wrong can be fatal in the game. Using that setup, they tested six frontier model snapshots, including Claude Opus 4.6, Claude Sonnet 4.6, GPT-5.1, and Gemini 3.1 Pro. The four larger models slashed their verification rates by 60-85% when paired with a consistently accurate partner. Two smaller unnamed models showed almost no adjustment. Trust formation made agents faster and more profitable in the game - but errors reversed the discount sharply, and failures bunched close together kept models suspicious far longer than the same number of spread-out mistakes.

The paper arrives at a useful moment: multi-agent AI products are shipping faster than anyone has a framework for governing them. The finding that over-verification correlates with indecision rather than safety is the kind of result that should give pause to the "more skepticism is always safer" camp. Calibration, the authors argue, is the real governance target - not reflexive distrust.

The asymmetry between formation and recovery is the wrinkle worth watching: agents warm up to teammates faster than they cool down after being burned. In benign environments that looks like efficiency; in adversarial ones, where a compromised agent feeds bad answers strategically, it looks like an attack surface the paper does not yet map.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →