AI/ robotics · vision-language-models · ai-research

Gondola Teaches Robots to Point Before They Act

A new research model plans robot manipulation by grounding each step in pixel-level object masks instead of raw action prediction, boosting benchmark results.

Researchers have built a robot planning model that points at exactly what it means before it moves a single joint.

Gondola is a vision-language planning system that sits between a robot's eyes and its motors. Instead of mapping camera images and text commands straight to motor actions, it first produces a structured plan: a text instruction paired with pixel-level masks across multiple camera views, marking the exact object and location it's referring to. The team trained it on synthetic data built specifically for short-horizon grounded planning, referring to objects across different viewpoints, and chaining together longer multi-step tasks. Paired with a separate 3D execution policy that turns those plans into motion, the system posted state-of-the-art results on the GemBench benchmark and showed early signs of transferring to real robots.

Robot arms have a well-documented habit of grabbing the wrong object or losing the thread on multi-step tasks, partly because most end-to-end vision-to-action models give engineers no way to see what the model thinks it's targeting. Forcing the system to mark its target with a visible mask before acting creates a debuggable checkpoint, and the paper's ablation tests show that this pixel-level grounding, not just more training data, drives the performance gains.

It's a sign that the just-predict-actions-directly approach to robot learning has limits, and that making robots show their work may matter as much as making them bigger.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →