AI researchers have a new fix for vision models that invent objects that aren't in the picture.
A method called ResOT repairs that hallucination problem at inference time, without retraining the underlying model. Large vision-language models often hallucinate because the internal representations that drive hallucinations are tangled up with representations carrying genuinely useful visual information, so simply suppressing them can dull the model's broader abilities. ResOT instead isolates a narrow residual subspace where hallucinated and faithful representations diverge, then uses a technique called Gaussian optimal transport to calculate how each token's representation should shift toward the faithful pattern, adjusting the distance case by case. Tested on three large vision-language models, ResOT reduced object hallucination while also improving caption quality and multimodal performance across several benchmarks, according to the researchers.
Most hallucination fixes trade accuracy for caution: suppress the suspect signal and the model gets safer but duller. ResOT's authors report the opposite tradeoff, improving both reliability and performance, which would matter for anyone shipping captioning, accessibility, or document-analysis tools built on vision-language models that are too expensive to retrain on demand.
The paper does not name the three models it tested, and the code is only promised, not released, so treat the results as promising until someone outside the lab can reproduce them.