The vision system inside a medical AI model turns out to be smarter than the model itself.
The paper, 'How Medical VLMs Underutilize Their Vision Encoders: A Dermatology Perspective' (arXiv:2609.36557, posted September 30, 2026, not yet peer-reviewed), compares two components: the MedSigLIP vision encoder and the MedGemma vision-language model built on top of it. Using dermatology as the test case, the researchers found MedSigLIP's raw image classification beat MedGemma's full diagnostic output by an average of 10.26 percentage points, even with zero labeled examples for the target task. Few-shot linear probing - a quick check of how useful an encoder's internal representations are - backed up the finding. In other words, the eye works fine. It's what happens after the eye reports back that goes wrong.
That gap matters because these tools are pitched for real diagnostic support, and a system that sounds confident while ignoring its own best evidence is a worse failure mode than one that's simply less accurate. The paper's authors trace part of the problem to attention: a simple describe-then-decide prompting step, where the model narrates the image before diagnosing, raised its attention to visual detail by 30-40% during generation. Fine-tuning the model specifically for dermatology improved classification but made it worse at general medical question-answering elsewhere, a trade-off worth flagging for anyone deploying one fine-tuned model across multiple departments.
The fix the authors propose - frozen models, label-free prompting, plus a lightweight reranking step powered by the encoder itself - is a patch, not a redesign, and a reminder that bigger multimodal models don't automatically make better use of what they see.