Medical AI models that answer questions about scans or write up reports are increasingly fine-tuned using preference optimization, a technique that scores whole answers as better or worse. New research argues that approach is too blunt for medicine, where correctness usually comes down to a handful of decisive words.
Researchers examined Direct Preference Optimization (DPO), the standard method for this kind of fine-tuning, and found it puts the least reward weight on exactly the phrases that determine clinical correctness, things like which side of the body a finding is on or a specific lesion attribute. They also found that a common fix, swapping in human-written reference answers as the preferred response, backfires: models learn to exploit stylistic differences between the reference and generated text rather than medical substance. That shortcut raises preference scores on paper without making the model more medically accurate, a classic case of reward hacking. To address it, the researchers built a new objective called FiRe-MPO, which edits a model's own answers in minimal, targeted ways so only the clinically decisive phrases differ between a preferred and rejected response, and pairs each example with a version of the image that has had the relevant visual evidence removed.
This is a useful reminder that benchmark gains do not equal clinical gains: a model can look like it is improving on preference metrics while quietly learning the wrong lesson. For medical tools, where plausible but wrong answers carry real risk, catching that gap before deployment matters more than in most AI applications.
In tests on medical visual question answering and report generation, across one medical-specialist and one general-purpose vision-language model, FiRe-MPO beat both DPO and other fine-grained alternatives, and did a better job keeping the model's reasoning tied to the actual image.