Passing a medical licensing exam and reasoning like a doctor are not the same thing, and a new review argues our AI benchmarks mostly measure the former.
A structured narrative review maps three separate literatures: medical education assessment tools, clinical large language model benchmarks published since 2023, and general-domain methods for scoring long-form AI text. The authors judge existing instruments against six dimensions of clinical reasoning, including problem representation, temporal synthesis, differential and management reasoning, counterfactual reasoning, calibrated uncertainty, and faithfulness, and find that no single tool covers all six. Problem representation and differential or management reasoning are reasonably well tested, and two newer tools, TIMER-Eval and ER-Reason, target temporal synthesis and sequential diagnostic belief updating. Uncertainty and counterfactual reasoning evaluations are still emerging, and faithfulness is the weakest dimension, backed by just one causal-ablation study that only covered multiple-choice questions.
That gap matters because exam-style accuracy is not the same as trustworthy reasoning: a model can land on the right diagnosis while quietly ignoring half the patient's history, and most current scorecards would not catch it. As hospitals pilot LLM copilots for chart review and triage, the review's finding that faithfulness checks barely exist is the uncomfortable one: it means we mostly cannot verify whether a model's stated reasoning matches what actually drove its answer.
The paper does not build a better test, either; it offers a design rationale for one, which is a polite way of saying the hard part is still homework.