A new benchmark shows that letting AI agents share too much memory backfires almost as often as it helps.
According to a preprint titled "CoMemBench: Benchmarking Collaborative Memory Boundaries across Multi-Agent Workflow Topologies," posted to arXiv as paper 2609.32192 (a replacement version dated September 30, 2026), researchers built 800 composite multi-agent workflows spanning four task domains. Each workflow is built from a dependency graph that specifies which intermediate outputs are relevant to which downstream agents, and for how long they stay valid. CoMemBench scores five things: whether the workflow finishes, whether each step is independently verified, whether required handoffs actually reach the right agent, whether stale or irrelevant information stays isolated, and how many tokens the whole run costs. The paper is a preprint and has not been through peer review.
The authors' headline finding is a trade-off: giving agents broader shared context improves how much useful information is available, but it weakens isolation, letting stale or unverified data leak into places it shouldn't reach. They also report that which multi-agent system ranks best shifts depending on the workflow's shape and how much artifact contamination occurs, which undercuts any claim that one architecture is a safe default.
It is a useful reality check on "just give every agent full context" - that's a design trade-off, not a free upgrade, and until this preprint there was no standard way to measure what it costs.