A new robot control model has found a way to have its imagination and its speed too.
Researchers built GlanceWAM, a world-action model that splits imagining from acting. Most systems like this generate video frames to predict what a robot's actions will look like, but doing that fast enough to control a robot in real time has meant either slowing everything down or skipping the imagining step and losing accuracy. GlanceWAM dodges that by running a slow background process that imagines a single frame seconds ahead, while a separate fast process decodes actual robot commands every 48 milliseconds, untouched by the slower process. On a 24-task kitchen benchmark called RoboCasa, it hit 72.2% success versus 67.1% for a synchronous rival called Cosmos Policy, and it scored 99.0% on a separate benchmark called LIBERO, all while cutting control latency 24 times over. In real-robot tests with one and two arms, it beat a model called pi-0.5 without needing any robot-specific pretraining data.
This matters because the speed-versus-accuracy trade-off has been the core bottleneck keeping video-based world models out of real robots. Decoupling imagination from control, rather than making imagination faster outright, is a genuinely different fix, and the latency numbers suggest it is a workable one.
Still, these are lab benchmarks run on a single A100 GPU, not a warehouse floor. Whether the lookahead frame stays useful outside curated demo tasks is the question that decides if this is a real shift or just a clever benchmark win.