Medical AI models sometimes abandon a correct diagnosis the moment a sentence in the chart contradicts the scan, and researchers now know exactly which attention heads are responsible.
A new paper introduces CRAFT (Causal Responsibility and Failure Tracing), a method for pinpointing the specific attention heads inside medical vision language models that cause two distinct failure types. One, called arbitration failure, happens when misleading text overrides a correct visual diagnosis. The other, brake failure, happens when a model commits to a confident answer even though the image doesn't support it. The researchers found these failures come from separate, non-overlapping groups of attention heads, then tested the theory by surgically removing each group and watching the behavior change.
That separation matters because it means the fixes don't have to be the same. Disabling the arbitration heads cut down on text overriding images without hurting performance on normal cases, while disabling the brake heads made models appropriately abstain when the visual evidence was weak. Both interventions worked across multiple medical visual question answering benchmarks and model architectures, without retraining.
It's a narrow fix for a narrow problem, not a cure for hallucination generally, but it's the kind of head-level accountability regulators will eventually want before these models get near a real patient chart.