Researchers built a benchmark that breaks a popular class of AI planning models, then found a fix that mostly repairs it.
The benchmark, called SLIM, asks a simulated robot arm to push small objects around a tabletop toward goals given either as an image or as a sentence. A world model called LeWM, which had aced a similar benchmark called PushT, solved under 1% of SLIM's trials, while a scripted controller with direct access to the simulator's internal state solved every tier. Diagnostic probes traced the failure to the model's encoder: its internal representation barely changed when the robot acted, and researchers couldn't even decode the pusher's or the objects' positions from it. Adding one extra training signal, a loss that forces the encoder to predict actions from pairs of latent states, fixed the probes and pushed success from 0.3% to 35%.
That's a quiet admission that a chunk of recent world model research may work by accident, succeeding only when enough of the frame happens to move for the model's shortcuts to still function. The paper's action-sensitivity probe, which needs no simulator access, offers a cheap pre-flight check before anyone trusts a world model's plans. That matters anywhere a company proposes letting a model simulate outcomes before a robot, or an agent, acts in the real world.
On the repaired model, a language-goal head reached 84% success on navigation tasks, versus 100% for the version given the goal as an image. Precision pushing is harder: a single sentence describing a goal rarely finishes a push, and even breaking the task into a sequence of stage-by-stage sentences only lifts success on the medium and hard pushing tiers from 4% to 25%. Useful, but a reminder that telling a robot what you want is still a weaker interface than showing it.