Researchers have found a way to make AI translation models actually pay attention to the images meant to disambiguate their text, and it works significantly better than existing fixes.
The technique, called Metric-based Loss Weighting, increases the training penalty for words whose correct translation depends on an accompanying image. To find those words, the researchers compare a model's output probabilities with and without visual input, using a measure called Point-wise Cross-mutual Information, or PCXMI. They also built a second, congruency-based version of that metric, and found combining both approaches worked best. The team fine-tuned three pretrained multimodal LLMs across three language directions to test the method on image-guided translation.
Multimodal translation systems have long had a habit of accepting images as input while quietly ignoring them, since text alone is often enough to produce a plausible translation. That makes benchmarks built specifically to test visual grounding, like the CoMMuTE contrastive dataset used here, more revealing than standard translation scores. On CoMMuTE, the new method beat standard fine-tuning by more than 7 percentage points in accuracy, without dragging down general translation quality.
It is a narrow fix for a narrow failure mode, but it is a useful reminder that giving a model extra data is not the same as making it use that data.