AI/ vision-language-models · spatial-reasoning · ai-training · computer-vision

Study Teaches AI to See Depth Without the 3D Hardware

A new training method called GPD improves vision language models' spatial reasoning by showing 3D geometry only during training, not at runtime.

Researchers have found a way to make AI vision systems better at judging depth and direction without ever running 3D sensors at inference time.

The method, called Geometry-Privileged Distillation, tackles a known weak spot in vision-language models: they take in flat RGB images and struggle to reason about real-world space, like how far apart two objects are or which way a hallway turns. Prior fixes either bolt 3D processing onto the model when it runs, which is slow and expensive, or train on reward signals that only check whether the final answer is right, ignoring where the reasoning went wrong. GPD instead feeds the training process depth maps, semantic labels, and bird's-eye-view data as text, but only to a teacher model and only when the student got the answer wrong. The deployed model never sees any of that extra data; it stays RGB-only.

That distinction matters because it sidesteps the usual tradeoff between accuracy and deployment cost. Most spatial-reasoning fixes either slow the model down with added 3D inputs or improve answers without improving the underlying perception. GPD's bet is that correcting the model only on its mistakes, using ground-truth geometry as a tutor rather than a crutch, teaches better spatial judgment without the runtime tax.

On a 4-billion-parameter model, GPD scored 57.1 on VSI-Bench and beat both standard reinforcement learning and an answer-only variant across four other spatial benchmarks. Whether that translates to robots or AR systems that need genuine real-time spatial awareness, rather than just better benchmark scores, is the harder question the paper doesn't answer.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →