AI/ ai · medical-imaging · dataset-contamination · research-integrity

Brain-Tumor MRI AI Benchmarks Are Quietly Leaking Data

A new audit finds popular brain-tumor MRI datasets riddled with training-test overlap, undermining AI models' near-perfect accuracy scores.

The brain-tumor MRI datasets behind years of near-perfect AI accuracy claims are full of leaks.

A new audit examined the three most widely used public brain-tumor MRI corpora and checked them for three layers of contamination: duplicate images, patient overlap, and source-label leakage. The dominant corpus turned out to have a near-twin of 28.8% of its official test images sitting in its own training split. A second corpus leaks 22.3% of its test images as byte-identical copies of training images, and 95.5% of traceable test images share a patient with the training set. The researchers also found that file-header metadata alone, with no visible anatomy, could separate tumor from non-tumor scans at 0.959 balanced accuracy - roughly matching fine-tuned ResNet models trained on the actual images.

The unsettling finding is that removing every identified leaked image barely changed reported accuracy. That means a stable leaderboard score, the thing researchers usually point to as proof a model works, says nothing about whether it learned to spot tumors or just learned to exploit dataset quirks. Years of published results built on these benchmarks now need a second look.

The team tested this across nine architectures and checked it against a chest-radiograph dataset as a negative control, so this isn't one model having a bad day. When an entire field reports 98%-plus accuracy on the same benchmarks, the more useful question isn't how it was achieved - it's whether the test was rigged by the data itself.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →