AI/ robotics · ai · computer-vision · world-models

ProWAM Teaches Robots to Plan With Sparse Visual Goals

A new robotics model predicts sparse visual sub-goals instead of full video, boosting real-world task success by 15 percentage points over the best baseline.

A new robot-control model skips frame-by-frame video prediction and still finds its way around rooms it has never seen before.

Researchers built ProWAM, a world action model that predicts a robot's next moves alongside a handful of sparse sub-goal images instead of generating a full video of the task. That sub-goal sequence gives the action system waypoints to aim for, and it is learned from large stockpiles of video that have no action labels attached. ProWAM runs its video model once per task to cache those sub-goal features, then relies on a lighter action-generation step during execution rather than repeatedly re-rendering video. In testing, it hit 85.8% on the LIBERO-Plus benchmark and 75.7% on a randomized RoboTwin suite, and in zero-shot real-world trials in new rooms it succeeded 70% of the time, up from 55% for the strongest prior model.

Full-video world models are slow because they have to imagine every frame between here and the goal, which makes them impractical for robots that need to react in real time. Sparse sub-goals cut that cost without losing the visual roadmap that helps a policy generalize to unfamiliar scenes, which is the actual bottleneck for warehouse and home robots today. The jump from 55% to 70% success in new environments is the number worth watching, since simulation benchmarks routinely flatter how well a model handles the unpredictable.

One asterisk: on the tougher RoboCasa365 benchmark, ProWAM only ranked 4th overall and managed 18.2% on its hardest unseen-task split, a reminder that sparse foresight helps but hasn't solved the long-horizon problem.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →