AI/ ai-agents · benchmarks · enterprise-software

ERPBench Shows AI Agents Fake Success on Enterprise Software

A new ERP benchmark finds AI agents can appear to succeed while saving wrong data to enterprise records, exposing a reliability gap.

A new benchmark says AI agents that ace desktop and web tasks fall apart the moment real business records are on the line.

Researchers built ERPBench, a benchmark that tests screenshot-only computer-use agents against a live, reproducible ERP system - the software that runs finance, procurement, inventory, and customer operations at most large organizations. Instead of grading whether an agent clicks the right buttons, ERPBench checks the actual values an agent writes into the database against ground truth. The team also built a harness that requires human approval before an agent can act, though ERPBench itself lets agents run autonomously for testing. Six closed and open-source agents were evaluated.

The results are rough: some agents saved a form in up to 85% of runs, but the value they wrote was correct in as few as 3% of those cases. That gap matters because ERP errors don't show up as a broken webpage - they show up months later as a wrong invoice, miscounted inventory, or a customer order shipped to the wrong address.

General-purpose GUI benchmarks have told a rosy story about agent competence; ERPBench is a reminder that saving a form and getting it right are two very different achievements.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →