A robot that follows one spoken instruction is old news. Getting four of them to follow different instructions without crashing into each other is a much harder problem.
A paper posted to arXiv this week (arXiv:2609.35965) formalizes multi-agent vision-and-language navigation as a coordination problem for the first time, rather than treating it as a solved side-case of single-agent navigation. The researchers built a benchmark called MAVLN: 11,724 episodes across 145 scenes, with teams of up to four agents working under three different instruction setups. Missions are broken into subtasks that carry dependency and resource constraints - one robot might need to hold a door while another passes through, for instance. Alongside the benchmark, the paper introduces TRISS, a baseline system that pairs an LLM-based scheduler with shared memory of what each agent has already explored, plus a mechanism for resolving collisions between agents' planned routes.
The interesting part isn't the benchmark itself - it's the admission baked into the results. The paper's own baseline, TRISS, leaves what the authors call 'substantial room for improvement' in scheduling, planning, and execution. That's notable because most navigation research still assumes a single robot and a single instruction, while multi-agent coordination is closer to how warehouse and delivery robotics actually operate.
Teaching one robot to follow directions was the easy part - teaching four of them to share the workspace without arguing is the actual product.