AI/ egocentric-video · video-language-models · benchmarks · wearable-ai

Benchmark Probes Whether Captions Can Give AI Video Memory

CapMem tests 33.7 hours of wearable-camera footage across 75 videos to see if text captions beat raw video question-answering for long-term recall.

A new benchmark says your smart glasses might remember your day better as text than as video.

Researchers built CapMem, a benchmark of 75 egocentric videos totaling 33.7 hours, paired with 1,000 multiple-choice questions across 16 everyday scenarios. The goal: test whether text captions can stand in for raw video when an AI assistant needs to recall something that happened earlier. On videos longer than 20 minutes, captioning the footage in 30-second or 60-second chunks and then answering questions from those captions beat feeding the raw video directly into a model, in 10 of 12 models tested at the 30-second window and 8 of 12 at the 60-second window. A retrieve-and-verify system that pulls relevant captions and double-checks them against the question pushed accuracy up by as much as 5.3 points.

Wearable AI assistants cannot simply feed hours of footage into a model. Vision-language systems cap how many frames they can process, and costs balloon with every added frame. Captioning turns video into cheap, searchable text that sidesteps the frame-budget problem entirely, which is a more practical route to 'the AI remembers your day' than waiting for models with ever-longer context windows.

It is also a quiet admission that today's long-context video models are not there yet: captions won not because text is inherently smarter than pixels, but because raw video question-answering still buckles once footage runs long.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →