A group of researchers has drafted a rubric for judging how well AI chatbots reason through medical cases, not just whether they land on the right diagnosis.
The framework, laid out in a new arXiv paper, pulls together ideas from medical education tools like OSCE exams and Key Feature Problems, existing clinical AI benchmarks such as MedR-Bench and HealthBench, and general LLM evaluation concepts like the Factuality-Validity-Coherence-Utility taxonomy. It scores free-text answers to standard clinical vignettes across several dimensions, including a "groundedness" measure adapted from that factuality taxonomy. The rubric also includes behavioral anchors, rules for when each criterion applies, and a separate flag for safety-critical errors specific to a given case. The authors are explicit that it has not yet been tested for inter-rater reliability, construct validity, or clinical utility.
Most AI labs tout their models' medical performance through benchmark scores that check whether an answer matches a reference diagnosis. This rubric targets something those benchmarks skip: whether the reasoning behind the answer actually holds up, which matters more to a clinician deciding whether to trust an AI's suggestion than whether it happened to land on the right answer. A scoring framework is only as useful as the humans applying it consistently, and that part remains unproven.
Call it a grading rubric published before anyone has checked whether two graders would mark the same paper the same way.