AI/ ai · medical-imaging · benchmarking · brain-mri

New Benchmark Reveals Brain MRI AI Scores Are Shakier Than Thought

A new protocol shows brain MRI anomaly detectors are often ranked by measurement quirks, not just model quality.

A new evaluation protocol finds that brain MRI anomaly-detection scores often hinge on technical choices nobody reports, not just model quality.

Researchers built MIRTO, a testing protocol that makes explicit the hidden decisions behind unsupervised anomaly-detection scores: how maps get aligned with reference images, how thresholds get set, and which false-positive budget gets used. They ran it on four methods trained on healthy brain scans and tested on 312 subjects from the BraTS 2020 dataset, repeating each comparison across 15,552 evaluation pipelines. One finding: a simple axis-order mismatch between stored maps and the reference image dropped a diffusion model's voxel-level AUROC from 0.873 to 0.583, while its slice-level score barely moved. Another: a Dice-score advantage that looked statistically significant under one threshold-setting method disappeared once false-positive rates were equalized across methods.

That second result is the one that matters most. If a leaderboard-topping result evaporates once you control for false-positive burden, the ranking was measuring the evaluation setup, not the model. A training-free tweak to one method's aggregation step raised its Dice score by 0.052 at equal false-positive burden - a bigger swing than some published innovations deliver.

Benchmark hygiene problems like this aren't new in machine learning, but medical imaging raises the stakes: these scores feed decisions about which models get trusted near real diagnoses. The researchers are upfront that their own protocol was developed and tested on the same cohort it evaluates, so they're calling their nine hypotheses exploratory, not proven. Worth remembering next time a paper claims state-of-the-art.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →