AI/ ai-agents · benchmarks · healthcare · finance

New Benchmark Finds AI Agents Flunk Real World Office Work

A new benchmark called DAYJOB gave AI agents real healthcare and finance tasks, and even the best model passed under a quarter of them.

AI agents still can't do a full day's work without supervision.

Researchers built DAYJOB, a benchmark of 130 open-ended tasks drawn from healthcare and finance professionals: 50 healthcare tasks and 80 finance tasks, each requiring a professional an estimated 13.6 to 16.6 hours to complete. Every task runs in a containerized environment and gets graded against an expert rubric with a median of 47.5 to 57.5 binary pass-fail criteria, and an attempt only counts as a pass if it clears every single one. The researchers tested 30 model configurations from 13 developers. The best performer, Claude Opus 5.5, passed just 24.7% of healthcare tasks and 23.9% of finance tasks, while the median configuration passed only 0.6% and 2.5%.

That gap between top and median performance matters more than the headline numbers. It suggests benchmarks built on short, well-specified prompts have been masking how badly agents handle the ambiguity of real professional work, like figuring out which documents are relevant or whether a request's premise even holds. The researchers found agents happily accepted false premises and carried bad inputs through otherwise coherent-looking analyses.

A model that aces a coding benchmark can still fail at the kind of multi-hour, loosely-specified work that fills most professionals' calendars.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →