A new benchmark shows most AI models still can't edit a chart without rewriting half of it.
Researchers released ChartRevise, a dataset and evaluation protocol for testing whether AI systems can make exact edits to chart-generating code. The dataset includes 92,438 editing records spanning 344 distinct edit types across 20 chart types and three plotting libraries, built systematically on the "grammar of graphics" framework. Unlike prior chart benchmarks that mostly check whether edited code runs or whether the resulting chart looks right, ChartRevise's evaluation separately scores whether a model completed the specific request, made unrelated changes it shouldn't have, or missed necessary follow-on updates elsewhere in the code. Testing five models against four existing benchmarks, the researchers found that fine-tuning on ChartRevise improved mean requirement recall by 16% and raised the "exact-edit success" rate by 22%.
Chart editing is a stand-in for a problem that shows up across code-assistant tools generally: can a model do the one thing it was asked, without touching anything else. That combination, finish the request but don't overreach, is exactly what today's broader coding benchmarks tend to blur together, measuring only whether code executes or output looks plausible.
Given how much enterprise interest is going into on-demand chart editing in BI tools, a 22% jump on a narrow, well-defined benchmark is a modest but concrete sign of where the real gaps still are: not whether code runs, but whether the model knows when to stop.