A new benchmark called CHI-Bench puts AI agents through real healthcare busywork, and most of them fail.
Researchers built CHI-Bench, a simulator spanning 20 healthcare apps and 87 tool interfaces, and tasked AI agents with three real workflows: provider prior authorization, payer utilization management, and care management. Each task forces an agent to juggle multiple roles, hand off work between them, hold multi-turn conversations like peer review calls, and follow a managed-care operations handbook made up of more than 1,290 documents. Across 30 combinations of agent harnesses and underlying models, the best performer resolved only 28.0% of tasks, and none cleared 20% when tested for strict, repeatable success. Asking an agent to do everything in one continuous session, rather than breaking the work into steps, dropped performance to 3.8%.
That gap matters because healthcare paperwork is exactly the kind of work agent vendors promise to automate first: rule-heavy, repetitive, and expensive to staff. If agents can't reliably navigate prior authorization today, pitches about AI fixing healthcare bureaucracy need a lot more evidence before anyone trusts them with real patient cases.
The researchers suspect the same failure pattern, dense rules, multiple roles, decisions that are hard to undo, will show up well beyond healthcare, in any enterprise workflow where mistakes are costly to reverse.