A new benchmark exposes how often video-search AI gets the right answer without ever really watching the video.
Retrieval-augmented generation, the trick of fetching relevant material before answering a query, has moved from text into long video, where the method is called VideoRAG. Researchers identified two problems with how the field evaluates it: existing benchmarks let models answer correctly without actually retrieving the right video evidence, and prior systems apply one fixed combination of modality and time granularity to every query, ignoring that different chunks need different approaches. To fix this, they built V-RAGBench, where each answer is tied to a single specific evidence chunk, and CARVE, a training-free method that runs several retrieval configurations in parallel and picks the winning one per chunk before generation. On the new benchmark, CARVE beat eight recent VideoRAG baselines on both retrieval and generation stages.
That matters because video RAG is quietly becoming the backend for tools that search security footage, lecture recordings, and personal video logs, and a system that is right for the wrong reasons will eventually fail somewhere expensive. By tying each answer to one verifiable chunk, V-RAGBench lets researchers tell whether a wrong answer came from retrieving the wrong moment or from misreading the right one, a distinction most prior benchmarks couldn't make. CARVE's gains held on both first-person footage and third-person footage, which suggests the fix isn't tuned to one video style.
Training-free is the selling point, but running parallel retrievers for every query is still more computation, not less. That's a reminder that fixing an evaluation gap and fixing a cost problem are two different projects.