AI/ ai-agents · benchmarks · arxiv · research

New Benchmark Pits AI Agents Against Human Experts

A new benchmark adds 400 expert tasks across law, finance, industry, healthcare, and science to test agents on real professional work, not exams.

A new benchmark wants to know if AI agents can handle real professional judgment calls, not just answer trivia.

Researchers have published $OneMillion-Bench ($OMB), a set of 400 expert-written tasks spanning law, finance, industry, healthcare, and natural science. Unlike typical benchmarks built from exam questions or tidy coding problems, these tasks ask an agent to pull authoritative sources, weigh conflicting evidence, apply domain-specific rules, and make a decision under real constraints. Grading uses a rubric that scores factual accuracy, logical coherence, practical feasibility, and professional compliance - so the reasoning path counts as much as the final answer. The paper, posted to arXiv on October 9, 2026, describes the benchmark's design and scoring protocol.

That focus on process over output matters because most agent benchmarks - think bar-exam questions or software-ticket tasks like SWE-bench - reward getting to a correct answer by any route. $OMB is closer to how a manager actually judges a junior analyst: not just what you concluded, but whether you checked the right sources and followed the right rules to get there. If it holds up, it could become a sharper gut-check for whether agents are ready for billable, liability-bearing work.

One catch: this paper only lays out the test. It does not report how any model or agent actually scored on it. The interesting number - how far language agents really are from human experts - is still unanswered.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →