A new benchmark calibrated to real finance work finds that AI agents can build spreadsheets. They just can't build the kind a finance team would actually use.
Researchers published MBABench, a benchmark evaluating large language model agents on complete financial spreadsheet tasks: financial modeling, forecasting, and scenario analysis constructed from scratch based on high-level instructions. Unlike previous spreadsheet benchmarks, which test question-answering or single-formula edits, MBABench grades agents across three dimensions: accuracy, formula correctness, and formatting quality. Those standards reflect how these documents are actually reviewed and revised by multiple stakeholders. The Claude family of models ranked highest overall and produced the most polished outputs in qualitative review.
The benchmark exposes a specific failure mode: agent performance degrades sharply once complexity extends beyond a few chained calculations, which is precisely where real financial modeling lives. For enterprise buyers evaluating AI-assisted finance tools, that cliff is a gap product demos are not designed to reveal.
Finance has historically been among the sectors most eager to adopt productivity tools and most burned when the edge cases arrive. That the strongest agents still fall short of professional standards here suggests current AI is closer to a capable intern than a reliable analyst.