AI systems that draft chest X-ray reports are being graded on a curve nobody talks about: the writing style of the human report they're compared to.
A new study tests nine AI radiology report generators against the MIMIC-CXR dataset using a standard scoring metric called RadCliQ-v1. Researchers built a tool called ReRef that rewrites the human reference reports, trimming shorthand, condensing routine findings, and changing formatting, without altering the actual medical meaning. When the same nine models were rescored against these reworded but clinically identical references, the leaderboard shuffled. One model, Libra, dropped from first place to second, while another, CheXOne, jumped from third to first, simply because the reference reports described normal findings more briefly.
That matters because hospitals and vendors use these leaderboards to decide which AI report writer is safe to deploy. If a benchmark score reflects how briefly a radiologist described normal findings rather than whether the AI caught the right clinical details, it is measuring prose habits, not diagnostic accuracy. The researchers argue current metrics fail to separate the two, rewarding or punishing an AI report for matching a reference's writing conventions rather than its clinical substance.
The team has released the underlying dataset, 120 radiologist checked pairs of original and rewritten reference reports, so future benchmarks can actually be tested for this blind spot instead of quietly baking in one hospital's paperwork habits.