Researchers have found a way to make AI vision models hallucinate less, without retraining them.
A new paper describes AIMS (Adaptive Information Multi-source Steering), a training-free framework for large vision-language models (LVLMs). The method tracks three sources of context during generation - the image itself, the text prompt fed in beforehand, and the text the model has already produced - and builds a compact profile for each. At each decoding step, AIMS checks how closely the model's current output aligns with each source and adjusts, layer by layer, how much weight that source gets. Tested across multiple LVLMs and decoding strategies, the approach reduced object hallucination while keeping general multimodal performance intact, according to the paper.
Most existing anti-hallucination fixes just crank up the model's attention to the image, on the assumption that ignoring the picture is the whole problem. This research suggests that's only part of the story: LVLMs already lean on visual input more than assumed, and text context - both the prompt and what the model has already said - also shapes whether it starts inventing objects. The fix, in other words, is balancing all three sources rather than amplifying one.
That balancing act is still inference-time plumbing, not a cure. It makes a model lie about what's in a photo less often. It doesn't make the model understand the photo.