AI/ robotics · world models · jepa · ai research

New Robot World Model Skips Pixel Reconstruction

A new world model for robots learns dynamics without reconstructing pixels, and planning in its policy's noise space beats sampling raw actions.

A new robotics paper trains an AI world model that never bothers reconstructing pixels, and it plans better because of it.

The system, called LeWAM, is a bidirectional transformer trained end-to-end on a JEPA-style latent space, the same self-supervised representation-learning idea associated with Meta's AI research group. Instead of rebuilding images frame by frame, it learns forward dynamics, backward dynamics, inverse dynamics, and a policy, all from one compressed representation. The researchers found that robot and object state are easier to read out of LeWAM's latent space than out of a standard forward-only JEPA world model, and far easier than out of a model built on pixel reconstruction. In closed-loop tests, a robot controlled by LeWAM matched the performance of one controlled by a separately trained flow-matching policy of the same size, while also coming with a usable model of the world.

World models are supposed to improve planning by letting software imagine outcomes before acting, but an imagined outcome only helps if the imagination is accurate. The paper's sharper finding is about where you search for a plan, not just how you represent the world: sampling raw actions lets a planner exploit the world model's own blind spots, while searching in the noise space of the policy network produced more reliable results.

It is a narrow fix, but it names a problem that quietly undermines a lot of model-based robotics work - plans that look flawless inside a glitchy simulated world and fall apart the instant a real robot has to act on them.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →