A new benchmark says AI agents still can't run a business.
Researchers behind a new arXiv preprint, posted September 30 and not yet peer-reviewed, introduce EnterpriseBench, a testing suite for large language model agents on enterprise strategic reasoning. It folds existing enterprise and financial QA datasets into one foundational suite, then adds three interactive scenarios: a management-consulting simulation where the agent diagnoses a client's problem through multi-turn questioning, a version of the classic Beer Game supply-chain simulation that tests inventory decisions under delayed feedback, and an "Enterprise Digital Twin" that simulates workforce, risk, and project planning. The authors ran nine agent methods across four backbone models, and none handled the full range of tasks reliably.
Most enterprise AI benchmarks measure whether a model can pull a number out of a filing or answer a finance trivia question - useful, but not the same as making a call on incomplete information and living with the consequences over several turns. EnterpriseBench targets that gap directly, and the fact that no method-model combination held up consistently is a useful reality check for anyone pitching autonomous "AI employees" this year.
It's a preprint, so treat the numbers as a first read rather than a verdict - but the test design itself is a fair admission that most agent benchmarks have been measuring the wrong thing.