AI/ ai · finance · benchmarks · enterprise

Finance Spreadsheet Benchmark Finds AI Agents Still Fall Short

MBABench tests AI agents on end-to-end financial modeling and finds performance drops sharply once task complexity stacks beyond a few steps.

A new benchmark calibrated to real finance work finds that AI agents can build spreadsheets. They just can't build the kind a finance team would actually use.

Researchers published MBABench, a benchmark evaluating large language model agents on complete financial spreadsheet tasks: financial modeling, forecasting, and scenario analysis constructed from scratch based on high-level instructions. Unlike previous spreadsheet benchmarks, which test question-answering or single-formula edits, MBABench grades agents across three dimensions: accuracy, formula correctness, and formatting quality. Those standards reflect how these documents are actually reviewed and revised by multiple stakeholders. The Claude family of models ranked highest overall and produced the most polished outputs in qualitative review.

The benchmark exposes a specific failure mode: agent performance degrades sharply once complexity extends beyond a few chained calculations, which is precisely where real financial modeling lives. For enterprise buyers evaluating AI-assisted finance tools, that cliff is a gap product demos are not designed to reveal.

Finance has historically been among the sectors most eager to adopt productivity tools and most burned when the edge cases arrive. That the strongest agents still fall short of professional standards here suggests current AI is closer to a capable intern than a reliable analyst.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →