A new technique squeezes video tokens from both directions at once, and vision-language models get noticeably faster without getting dumber.
Researchers describe VETO (Video Efficient Token Optimization), a plug-in that doesn't require retraining. It compresses tokens two ways: merging similar-looking patches within a single frame, then merging frames that look nearly identical to their neighbors. The trick is doing the spatial squeeze first, which makes the frame-to-frame comparison much cheaper to compute. Tested on LLaVA-OneVision-7B, VETO cut inference time by up to 45% while matching or beating baseline accuracy.
Long videos are exactly where today's VLMs choke, since token counts explode quadratically as clips get longer, making real deployment costly. VETO's edge shows up most under tight budgets: given only 10% of the usual token budget, it held 55.7% accuracy versus 54.9% for VFlowOpt, 52.6% for VisionZip, and 47.9% for FastV. That's not just incremental, it's the difference between a model being usable on a budget and falling apart.
Still, these are benchmark numbers from the paper's own tests across LLaVA-OneVision, InternVL-2.5, and LongVA, not independent validation, and 'drop-in plug-in' claims have a way of getting messier once they meet someone else's production stack.