A new robotics model treats picking an action and predicting what happens next as the same language problem.
Researchers built a system called Spatial Language Modeling that represents scene contours, goals, action targets, and future states with one shared vocabulary of coordinates and tokens. A single autoregressive Transformer learns to choose actions and to predict how the scene will change, both through the same next-token prediction objective. The team trained it from scratch: first on random, unstructured interactions with objects, then on expert demonstrations that pair action and state sequences. On the Push-T object-pushing task, the model posted competitive results in simulation and beat the real-robot policy baselines it was tested against on task success and target coverage.
Most manipulation systems split "what to do" and "what will happen" into separate models or stages. Merging both into one next-token objective, using a single vocabulary of spatial tokens, is the actual contribution - a simpler training recipe that appears to hold up better on a physical robot. Ablations support that: joint action-and-state training beat action-only training, and pretraining on random play pushed performance up further.
One pushing task with one robot is nowhere near proof of a general-purpose robot brain, and "competitive" in simulation is not the same as "winning" - the real story here is better transfer to hardware, not a simulation breakthrough.