Multi-agent AI systems don't usually fail because the AI is dumb. They fail because nobody taught the agents to take turns.
A new study pulled 22,848 closed GitHub issues from 21 open-source projects built around LLM-based multi-agent architectures, then filtered that down to 944 issues specifically about multi-agent problems. The most common complaint, by far, was orchestration and execution breaking down - agents mishandling task handoffs, workflows stalling, or steps running out of order. The leading causes were workflow design flaws, broken tool integrations, and memory mismanagement. When developers fixed these bugs, the most common solution was not a better model. It was reworking the workflow itself.
That lines up with a pattern showing up across the agent boom: stacking several language models together exposes failure modes that never show up when you run just one. The bottleneck isn't reasoning ability, it's the coordination layer - routing tasks, passing context between agents, and keeping track of what already happened. That's a software engineering problem, not a model-capability one, and no benchmark about raw intelligence is going to fix it.
The researchers' top recommendation was blunter than any product launch: optimize the workflow. Not the model. The workflow.