A new benchmark tests AI agents on enterprise software by inventing an entire company from scratch, records and all.
Researchers built the Era by Eon Benchmark around a fictional company generated from an industry, size, business model, and seed. One shared entity graph feeds simulators of Salesforce, Zendesk, Slack, Gong, and other workplace tools, plus a separate generator that builds each product's internal databases from those same entities. Because every fact traces back to one graph, the benchmark can compute an exact answer key instead of relying on human graders. Across 23 generated companies, an automated realism check pushed the average score from 61.8 to 97.0, with zero records flagged as synthetic.
That consistency is the point: you can't test agents on real customer data, and earlier synthetic datasets had no way to verify their own ground truth. When the researchers ran nine models through the same 33 questions three times each, accuracy ranged from 42.4% to 76.8% - and only three of 36 pairwise differences held up after statistical correction. That's arguably the bigger finding: agent leaderboards may be measuring noise as much as skill.
A fake company scoring 97 out of 100 on realism is still, by definition, not real - but it beats grading agents against rubrics nobody can independently check.