AI/ video-ai · computer-vision · efficient-ai · video-grounding

A stripped-down AI model that finds moments in video from text

A lightweight architecture pinpoints video segments matching text queries using under 90 million parameters, far less compute than prior zero-shot systems.

A new video-search model finds the exact moment and location a sentence describes, using a fraction of the usual computing power.

Researchers describe Pocket-STVG, a lightweight system for spatio-temporal video grounding: matching a natural-language query to a specific object moving through a specific stretch of video. Instead of training one large end-to-end model, the team combined existing lightweight components, including a MobileViCLIP-based video encoder, an MDETR-derived spatial encoder-decoder, and a shared text encoder. Timing is handled by either a small 1D U-Net or a basic thresholding step, letting the same setup work in both weakly supervised and zero-shot modes. The whole pipeline runs on fewer than 90 million parameters, and video features can be precomputed before any query ever arrives.

Most spatio-temporal video grounding systems lean on heavyweight multimodal large language models or complex training pipelines, which makes them costly to run at scale. Pocket-STVG matches weakly supervised baselines and beats earlier zero-shot approaches while using a fraction of the memory and compute those methods need. That efficiency matters more than incremental accuracy gains for anyone trying to search large video archives rather than demo a single clip.

It is a reminder that in video AI, bigger model and better model are not the same claim. Sometimes recombining smaller, proven parts wins on the metric that actually ships: cost per query.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →