A new paper argues that AI agent teams don't fail because the agents are dumb - they fail because nobody owns the shared facts.
Researchers studied multi-agent AI systems where every individual agent makes a locally valid decision, yet the group's combined output is wrong - what they call the global coherence problem. They prove what they dub the Observation-Aliasing Impossibility Theorem: if several possible versions of the world look identical to an agent but require different actions, that agent's best possible success rate caps at 1/k, where k is the number of indistinguishable versions. No amount of extra reasoning, messaging, or sampling raises that ceiling. They then propose a runtime layer that tracks shared state directly, using overlapping scopes and consistency checks, so agents propose actions but a separate harness approves and commits them. Across nine test setups, hiding the deciding fact from agents dropped scores from 40 out of 40 to 12-17 out of 40, close to the 1-in-3 chance baseline, and restoring that single fact brought performance straight back to 40 out of 40.
This directly undercuts the current pitch for multi-agent AI products, which often implies that stacking more agents, roles, or reasoning steps makes a system more reliable on its own. The paper's budget test makes the stakes concrete: ordinary agent teams blew through a shared budget in all 5 test runs, and even a visible live counter still left violations in 4 of 5 runs - only an explicit commit check got it to 0 of 5.
Translation: a multi-agent system is one state desync away from double-booking or overspending, and a smarter model won't fix that. Better bookkeeping will.