A new AI world model tackles long-horizon planning by refusing to cram short-term and long-term reasoning into a single latent space.
The system, called Dual-WM, comes from a preprint posted to arXiv on September 30, 2026 (arXiv:2609.37644). Rather than using one latent space to predict everything from the next frame to a goal 100 steps away, Dual-WM splits the work: a low-level model handles action-by-action transitions, while a high-level model plans using learned 'macro-actions' that span much longer stretches of time. The high-level model sketches latent subgoals, and the low-level model converts those into precise actions. The authors also introduce a training method called LoRe, which weights predictions differently depending on how far into the future they reach, to curb the way small errors compound during recursive rollouts.
Long-horizon planning is the known weak spot of latent world models: they're accurate at predicting what happens next but drift badly 50 or 100 steps out, as small errors snowball and distances within a single latent space stop meaningfully separating good goals from bad ones. Tested on five goal-conditioned visual control tasks, Dual-WM lifted mean success from 75.9% to 84.4% at a 50-step goal offset and from 61.4% to 69.5% at 100 steps, beating the strongest prior baseline, LeWM, by 30.8 percentage points at that longer horizon.
It's a narrow benchmark win, not proof of general intelligence. But splitting one overworked latent space into two specialized ones is a clean rebuttal to years of trying to make a single representation do both jobs, and the implementation is already public on GitHub for anyone who wants to stress-test it.