AI/ accessibility · multimodal-ai · data-visualization · ai-research

MLLMs Chart Descriptions Blur Data and Guesswork

A study of three multimodal AI models found they often can't tell readers which parts of a chart description are backed by data and which are speculation.

AI models that describe charts for accessibility often mix hard data with invented context, and they rarely flag which is which.

Researchers ran three multimodal large language models, including Gemini and GPT variants, through 102 visualizations pulled from four sources. They fed each chart under four different conditions, varying access to the image itself, accessible chart context like data tables, captions, and alt text, and prompts that withheld context. Across 1,224 resulting descriptions, they labeled each claim as DIRECT, DERIVED, or SPECULATIVE and checked whether the numbers models cited actually matched the source data. Giving models accessible chart context pushed Gemini and GPT toward more DIRECT, verifiable claims and improved numeric accuracy for some models, but adding the raw image on top of that context did not consistently help.

This matters because screen-reader users and other accessibility tools increasingly rely on AI-generated chart descriptions, and there's currently no reliable signal telling a reader when a model is reporting a data point versus making one up. The study found that a prompt section explicitly asking for real-world significance produced mostly speculative content anyway, and telling models to be cautious when context was withheld didn't make them noticeably more hedged.

It's the same trust problem that's dogged AI summarization for years, just relocated to a format where blind and low-vision users have fewer ways to independently double-check the output.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →