AI/ ai · benchmarks · chart-qa · multimodal-ai

A multi-agent system builds chart tests to find AI blind spots

A new agent harness turns sparse error descriptions into chart-reading tests that reliably expose where AI models actually fail.

A new system called ChartBmkAgent can turn a vague description of an AI model's weak spot into a working test for that weak spot, automatically.

Benchmark-building has become a bottleneck. Multimodal AI models improve fast; the tests used to probe them don't. Researchers built ChartBmkAgent to close that gap for chart-reading question-and-answer tasks, which require a model to both read a chart and reason about it. Instead of starting from templates or hand-collected data, a central "harness" coordinates several specialized agents to build full question-and-answer samples directly from a sparse description of an error category. Each agent has to produce evidence that its output still matches the original error category before the harness accepts it, and the system logs the reasoning behind every accept-or-reject call.

In testing, 300 generated samples produced model accuracies ranging from 32.7% to 84.3% depending on the error category, which the researchers take as a sign the samples are measuring real, distinct capability gaps rather than noise. More tellingly, when the system generated follow-up questions targeted at known weak spots in three different models, accuracy on those questions dropped to 50.0%, against 82.2% on matched control questions.

This matters because evaluators have historically struggled to keep pace with newly discovered model failures. A pipeline that converts a spotted weakness into a standardized test in the same workflow shortens that cycle considerably, and separate evaluator models confirmed that 86.4% of generated samples actually tested the category they were built for.

Worth noting: this is a single arXiv preprint, not yet peer-reviewed, and it leans on other AI models to grade whether the generated tests are valid - a self-referential setup that's useful but not proof the benchmarks will hold up outside the paper's own test bed.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →