AI/ robotics · vision-language-models · multi-robot-systems · ai-research

DuoMind Gives Robots a Shared Language for Teamwork

A new framework splits robot brains into a talkative planner and a precise doer, letting multiple robots coordinate on tasks without a central controller.

Researchers have built a way for teams of robots to talk to each other in plain concepts, not just raw coordinates.

The system, called DuoMind, gives each robot two brains. A vision-language-action model handles the fine motor work: gripping, moving, placing. A separate vision-language model sits above it, reading the task, the robot's surroundings, and messages from teammates, then issuing both low-level commands and semantic updates to other robots. There's no central dispatcher. Each robot reasons independently and shares what it learns. The researchers also built a new benchmark, RoboPoly, because existing tests weren't built for multi-robot, closed-loop coordination, and validated results on an existing one called RoboTwin.

Most robot AI progress has been single-robot progress: one arm, one camera, one task. Warehouses and factories need fleets that divide labor without constant human reassignment, and current approaches tend to fall back on rigid, pre-scripted handoffs rather than robots actually reasoning about what their peers are doing. Splitting perception-and-planning from execution, then letting the planning layer pass messages, is a more scalable path than building one monolithic model to do everything.

It's a research paper with a benchmark, not a product on a loading dock, so the real test is whether semantic chatter holds up when a robot's camera is wrong, a message is dropped, or ten robots are coordinating instead of two.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →