When AI agents represent competing users instead of one boss, teamwork makes things worse, not better.
Researchers tested five frontier models across 77 scenarios in four shared-resource environments: an API key setup where agents split a compute budget, a clinic where they share a calendar, a personal assistant setup where they share a group order or booking, and a merge queue where they share a release cutoff. In each case, they compared a single coordinator agent serving every user against a team of agents, each representing one user, with and without a channel to talk to each other. The teams did worse than the coordinator in every environment. Without a communication channel, coordination collapsed completely in two of the four environments, and in the personal assistant setup, the lone coordinator fulfilled a targeted user request about twice as often as the team did.
This matters because the whole pitch for personal AI agents assumes they can negotiate on our behalf when they bump into other people's agents over a calendar slot, a budget, or a deadline. This research suggests that assumption is shaky: agents stalled as teams grew, overrode each other's actions, and in some cases fabricated claims to get their way.
The proposed fixes - a team lead, explicit procedural instructions, a check that forces an agent to read its peers' messages before acting - sound less like clever AI design and more like basic office management, now applied to software.