AI/ ai agents · benchmarks · llm evaluation · enterprise ai

New Benchmark Shows AI Agents Struggle With Hidden Facts

A new benchmark called Era by Eon finds the best AI agent gets 18 of 24 hidden-fact questions right (75%), while weaker models manage just 6 of 24 (25%).

A new benchmark suggests enterprise AI agents are decent at math and bad at detective work.

Era by Eon started as a 27-question test where a generated company's data feeds an agent that must apply stated rules and compute an answer. When agents could run code, the four best models each landed 22 to 25 out of 27 - a gap too narrow to mean much. So the researchers added eight new question types built on facts nobody states outright. The records that look like they hold the answer actually point somewhere else, and the real explanation has to be pieced together from other data - in one example, the sales system logs a lost deal as a timing issue, but a recorded customer call blames a service outage instead. Across 12 agents tested three times per question (24 attempts each), the top agent got 18 right, a 75% score, while four of the six underlying models scored 6 or fewer no matter which agent program ran them, just 25%.

That gap matters because most business software doesn't state its facts cleanly. Tickets, transcripts, and CRM notes routinely contradict each other, and figuring out which one is lying is closer to what an analyst actually does all day than running a formula. A benchmark that rewards noticing a call transcript contradicts a sales log is testing judgment, not arithmetic, and judgment is where these agents are shakiest.

The category that broke everyone: picking the correct record out of several similar ones, like identifying which of three renewal offers a customer actually signed. All 12 agents combined got that right in only 1 of 84 tries - a hit rate that would embarrass a coin flip.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →