Give a team of AI agents a shared memory, and a single lie can spread through it almost unchecked, a new study finds.
Researchers built a test called the Correlated Promotion Benchmark to check whether AI systems can tell a genuinely new claim from one that has just been copied or reworded from something already in memory. One version, CPB-Static, uses a frozen set of pre-labeled examples; the other, CPB-Live, watches multiple AI agents write to and read from a real shared memory store while tracking where every claim originally came from. A separate "consumer" agent then answers questions using only what is in that shared store. The team tested eight different rules for deciding what gets admitted to memory, across four different AI model families.
The results are not encouraging. Policies that filter out duplicate claims end up throwing out plenty of true ones too, while policies that prioritize complete answers let in almost as many false claims as having no filter at all. The only approach that meaningfully cut false claims, checking whether a claim's declared source type was legitimate, still let through 6-9% of falsehoods, and once a false claim got into memory uncontested, agents repeated it as fact in 97-99% of checks.
For anyone building multi-agent systems on the promise of shared memory, that is a design flaw hiding behind a demo, not an edge case.