Researchers have built a way for teams of robots to talk to each other in plain concepts, not just raw coordinates.
The system, called DuoMind, gives each robot two brains. A vision-language-action model handles the fine motor work: gripping, moving, placing. A separate vision-language model sits above it, reading the task, the robot's surroundings, and messages from teammates, then issuing both low-level commands and semantic updates to other robots. There's no central dispatcher. Each robot reasons independently and shares what it learns. The researchers also built a new benchmark, RoboPoly, because existing tests weren't built for multi-robot, closed-loop coordination, and validated results on an existing one called RoboTwin.
Most robot AI progress has been single-robot progress: one arm, one camera, one task. Warehouses and factories need fleets that divide labor without constant human reassignment, and current approaches tend to fall back on rigid, pre-scripted handoffs rather than robots actually reasoning about what their peers are doing. Splitting perception-and-planning from execution, then letting the planning layer pass messages, is a more scalable path than building one monolithic model to do everything.
It's a research paper with a benchmark, not a product on a loading dock, so the real test is whether semantic chatter holds up when a robot's camera is wrong, a message is dropped, or ten robots are coordinating instead of two.