A new benchmark shows AI models get worse at reading charts as more charts get crammed into one image.
Researchers built ChartDensity-Bench to test how well multimodal large language models (MLLMs) reconstruct the actual numbers behind scientific charts. Rather than asking a model to describe a chart or answer questions about it, the benchmark asks for the raw underlying data, then scores the answer on structural reliability, completeness, parseability, and numerical accuracy. The test varies how crowded a single image gets, showing one, three, six, or nine charts at once. The team ran five recent MLLMs through all four density levels.
The results are not flattering. Reconstruction accuracy drops as chart count climbs, and the size of that drop varies a lot by model, meaning some systems degrade gracefully while others fall apart. Most tellingly, the exact same chart produces worse data extraction when it is surrounded by eight others than when it stands alone, which points to visual clutter itself as the culprit, not the chart's inherent difficulty.
That matters because real scientific figures rarely show up as single, isolated charts; they come in grids and panels, and this is the first benchmark to put a number on what that crowding actually costs AI readers.