Researchers have built a robot planning model that points at exactly what it means before it moves a single joint.
Gondola is a vision-language planning system that sits between a robot's eyes and its motors. Instead of mapping camera images and text commands straight to motor actions, it first produces a structured plan: a text instruction paired with pixel-level masks across multiple camera views, marking the exact object and location it's referring to. The team trained it on synthetic data built specifically for short-horizon grounded planning, referring to objects across different viewpoints, and chaining together longer multi-step tasks. Paired with a separate 3D execution policy that turns those plans into motion, the system posted state-of-the-art results on the GemBench benchmark and showed early signs of transferring to real robots.
Robot arms have a well-documented habit of grabbing the wrong object or losing the thread on multi-step tasks, partly because most end-to-end vision-to-action models give engineers no way to see what the model thinks it's targeting. Forcing the system to mark its target with a visible mask before acting creates a debuggable checkpoint, and the paper's ablation tests show that this pixel-level grounding, not just more training data, drives the performance gains.
It's a sign that the just-predict-actions-directly approach to robot learning has limits, and that making robots show their work may matter as much as making them bigger.