AI/ ai · multimodal-ai · robotics · autonomous-driving

Researchers Cut AI Vision Model Overhead With Token Trick

A new decoding method called DVD compresses 2D and 3D perception data into compact tokens to cut latency in multimodal AI systems.

A new technique called Dynamic Vector Decoding, or DVD, trims the token overhead that multimodal AI models rack up when they try to describe where objects are in an image or a 3D scene.

Researchers behind DVD built a way to convert bounding boxes, masks, and 3D bounding boxes into compact 1D vector sequences, then map those into discrete tokens a multimodal large language model can read and write directly. A lightweight de-tokenizer converts the model's output tokens back into the original 2D or 3D coordinates. The team tested DVD on five benchmarks - RefCOCO, SUN-RGBD, KITTI, Hypersim, and nuScenes - covering both flat image tasks and full 3D scene understanding. Across those tests, DVD matched or beat existing approaches while cutting token counts and inference latency.

That efficiency gap is the real story here. Today's multimodal models mostly describe object locations by writing coordinates out as plain text, which burns tokens fast, or by quantizing them into a fixed numeric range, which breaks down in 3D space where distances aren't bounded the way pixel coordinates are. Robotics and self-driving systems need both speed and precision at a scale no fixed-range scheme easily supports, so a method that handles 2D and 3D with one unified token scheme closes a real gap.

Still, a benchmark win on RefCOCO or KITTI is a long way from a car or a warehouse robot running this in real time. Plenty of efficient perception tricks have looked great in papers and never made it past the lab.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →