Turning a messy spreadsheet into a clean chart is still a brutal test for AI.
A new benchmark called Chart2Code puts large multimodal models through three escalating tasks: reproducing a chart from an image and a prompt, editing an existing chart by changing types or adding elements, and generating a chart from a long, dense table. The benchmark covers 2,023 tasks across 22 chart types and scores models on both whether the generated code runs correctly and whether the resulting chart actually looks right. Researchers tested 25 models, including GPT-5, Qwen2.5-VL, InternVL3/3.5, MiMo-VL, and Seed-1.6-VL. On the editing tasks, GPT-5, the best performer, averaged just 0.57 on code correctness and 0.22 on visual fidelity, out of a possible 1.0.
That gap between code that runs and a chart that looks right is the real story. It means models can produce plausible Python or JavaScript that executes fine while still misreading what a reference chart or a reworded instruction actually asked for. For any team hoping to automate dashboard or report generation, that's the difference between a tool you can trust and one you have to double-check anyway.
Benchmarks like this tend to get gamed within a year as labs train specifically to beat them. Worth watching whether chart-editing scores climb as fast as the much-hyped coding and math leaderboards have.