A new robot planning system skips the step everyone else needs: a picture of the goal.
Researchers describe the Grounded World Model, or GWM, a latent world model that plans robot actions directly from a text instruction instead of a goal image. Give it a candidate sequence of actions and a current camera view, and GWM predicts what the scene will look like next inside the visual space of a pretrained video-language embedding model. That same frozen embedding model scores how well the imagined future matches the task description, using negative cosine similarity as the cost function. Crucially, GWM trains only on offline, task-agnostic video-action pairs, with no labeled language instructions required, and the planner simply picks whichever candidate action scores lowest cost. On the WISER simulated benchmark, that approach solved 87% of 288 tasks involving unseen instructions and visuals, while ten separately fine-tuned vision-language-action models averaged just 22%. Scaled up with real robot data, GWM completed all 70 trials across 14 tasks requiring reasoning about ambiguous referring expressions in IsaacSim, matching a modular planner built around a frontier vision-language model, while the pi0.5 model managed 37 of 70. On an actual Franka arm, the same setup finished 55 of 60 pick-and-place sub-tasks, close to the modular planner's 52 of 60, all running locally on one consumer GPU.
The headline number here isn't accuracy, it's generalization. Most world models, including DINO-WM and LeWM, need a target image before they can plan, which is awkward for any task nobody has staged before. GWM's language-only goal setting, with zero extra fine-tuning, suggests embedding-matching could be a cheaper substitute for training a dedicated action model per task.
Still, 70 simulated trials and 60 real pick-and-place attempts is a small sample for a method this ambitious, and pick-and-place remains one of robotics' easier party tricks.