A new benchmark called UltraG-Bench finds that AI models good at describing ultrasound scans are bad at showing their work.
Researchers built the benchmark by annotating 40 public ultrasound segmentation datasets spanning 13 anatomical categories. It tests three tasks: instruction-guided segmentation, evidence-grounded visual question answering, and evidence-grounded report generation, adding up to more than 1.1 million annotations across the three tasks. The team evaluated 14 state-of-the-art vision-language models and found a substantial gap between what the models say and where they can actually point. They also built a companion system, UltraG-Agent, which pairs a vision-language model's reasoning with UltraSAM3's ultrasound-specific segmentation and improves both the predictions and the pixel-level grounding.
This matters because a model that says "this looks abnormal" without marking exactly which pixels it means isn't clinically useful, and it's hard to trust. Ultrasound is noisier and more operator-dependent than CT or MRI, which makes this gap especially relevant as hospitals start piloting AI-assisted diagnostic tools built on these same vision-language models.
An AI that can narrate a scan fluently is not the same thing as one you'd let circle the tumor.