A new technique trims the computational fat off robots' world-prediction AI, making it run up to twice as fast without any retraining.
Robots increasingly run on "world-action models," systems that predict their next camera frame and their next physical move at the same time, borrowing from video-generating AI to help the robot generalize to new situations. That dual prediction is slow, because the model has to reprocess every pixel-token of the imagined future frame at each step of denoising, the iterative noise-to-image cleanup process diffusion models rely on. Researchers built Sparse-WAM, a plug-in method that checks which parts of the predicted frame the robot's chosen action is actually paying attention to, then drops the rest instead of treating every token equally. Because that attention pattern barely shifts between consecutive steps, the team also built a lightweight engine called Pilot to reuse those relevance scores cheaply, avoiding overhead that would otherwise cancel out the savings. On the LIBERO benchmark and on RoboLab-120 with Cosmos 3 Edge, running on an RTX 4090, it delivered roughly 2.0x and 1.8x speedups while largely preserving task performance.
The real story here is what gets pruned. Earlier token-pruning methods for video diffusion models cut tokens based on visual fidelity, keeping whatever looked sharpest. Sparse-WAM instead asks which pixels the robot's action actually needs, a reframing that treats prediction as a means to control rather than an end in itself. For robots meant to run outside a data center, that distinction is the difference between a demo and a deployable system.
Translation: same robot performance, close to half the compute bill, so long as the trick holds up past two lab benchmarks.