Put multiple AI coding agents to work on the same codebase and they will quietly wreck each other's changes - a new benchmark shows the fix is scheduling the work up front, not watching for collisions after the fact.
Researchers tested parallel coding agents using NP-Bench, a three-arm benchmark that checks whether agents' merged code actually works, verified against real git merges rather than each agent's own passing tests. One setup gave agents no coordination, a second reacted to conflicts after they happened, and a third - built into a system called Nerveplane - planned ahead: it split work into separate scopes and ordered merges by who produces code a teammate consumes. The planner took clean, working integrations from 1 of 9 test scenarios to 9 of 9, and cut merge conflicts from 13 to zero. The advantage grew as more agents were added to the mix.
The scheduling approach also rescued a scenario the other two methods failed on every single time: a teammate silently changing a shared contract mid-task. With planning, a frontier model recovered cleanly 100% of the time and a smaller model 60% of the time, compared to zero for both reactive detection and no coordination. The gains held whether the underlying model was strong or weak, which suggests the real bottleneck in multi-agent coding isn't how smart the model is - it's how the work gets divided up.
The paper is refreshingly honest about its limits too: feeding agents more context facts didn't help long-context accuracy once the work already fit the context window - a reminder that more information isn't the same as a plan.