AI/ world models · robotics · language models · planning ai

New World Model LAGO Breaks Instructions Into Waypoints

Researchers built an AI planner that predicts its own subgoals from language, more than doubling success on tricky navigation tasks in tests.

A new AI planning system skips the usual tradeoff between picture-perfect goals and vague words, using the same model for both.

The team's system, called LAGO, is a single world model that does two jobs with one shared latent space: it predicts what will happen if the agent takes an action, and it translates a language instruction into a string of intermediate subgoals inside that same space. Instead of forcing the agent to hit each subgoal exactly, a soft-minimum alignment cost just rewards getting close, and the subgoals get recalculated as the agent moves. In tests spanning navigation and manipulation tasks, LAGO beat both flat and hierarchical planners that rely on image-based goals. The biggest jump came on long, curved paths with hazards to avoid, where it more than doubled the success rate of a goal-conditioned behavior-cloning policy trained on identical demonstrations.

That matters because language is the easiest way to tell a robot or agent what to do, but it has historically been the least reliable signal for actually planning a path - usually because teams bolt on a separate vision-language model to interpret it, and alignment drifts. LAGO's approach is grounding instructions in the planner's own learned space instead of outsourcing that translation, and it beat vision-language reward models at that job, including one given access to the real underlying dynamics.

It's a research result, not a product - long-horizon robot instructions have broken plenty of promising systems before, and real-world hazards aren't curated test environments.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →