A new benchmark says today's best AI video models still can't reliably explain what a data video is actually showing.
Researchers built DataVista, the first benchmark for "data videos" - the news-graphics style clips that mix animated charts with narration, common in business reporting and news coverage. It includes 961 real data videos and 6,775 questions split into three tiers: reading data straight off the screen, tracking how it changes over time, and understanding the narrative built around it. The team ran 19 mainstream multimodal large language models through those questions. Gemini-3.1-Pro came out on top but only managed 70.0% overall accuracy, well short of human expert performance.
The models performed worst on causal reasoning and narrative structure questions - the ones that ask why something happened in the data, not just what the numbers say. Feeding models more video frames or adding subtitles helped with basic data reading and tracking changes over time, but barely moved the needle on following the actual story. That gap matters for any product pitching AI as a tool to summarize earnings calls, news segments, or dashboards packaged as video.
A benchmark isn't a product, but it's a useful reminder that these models are still better at describing a chart than explaining what it means.