AI/ ai · benchmarks · multimodal-models · bias

Benchmark Shows AI Still Struggles to Follow Along With Video

A new benchmark finds leading AI models falter at pinpointing moments in audio and video, with accuracy gaps tied to who's speaking.

A new benchmark says today's AI video models are good at trivia but bad at timing.

Researchers at the Vector Institute released SONIC-O1, a benchmark built from 60 hours of real footage across 231 clips and 13 everyday conversational settings, with 4,958 human-verified annotations and demographic metadata attached to each clip. It tests multimodal large language models on three tasks: summarizing what happened, answering multiple-choice questions, and pinpointing exactly when something happened in the audio-video timeline with a supporting rationale. Across both closed-source and open-source models, multiple-choice accuracy was the closest contest. Temporal localization was not: the best closed-source model beat the best open-source model by 22.6%.

That gap matters because knowing what happened is a different skill than knowing when it happened, and the when is what live captioning, surveillance review, and video search tools actually need. The benchmark also found accuracy swings of up to 21.4% across demographic groups on the same temporal task, meaning the models' blind spots are not evenly distributed.

Benchmarks promising to be comprehensive show up constantly; what's notable here is which skill breaks first, not recognition, but timing.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →