Vision-language AI models often ace multimodal benchmarks while making claims about images that the images don't actually support.
Researchers studying reinforcement learning for large vision-language models found that before RL training, nearly 28% of Qwen2.5-VL-7B's correctly answered responses on four multimodal reasoning benchmarks contained at least one visual claim unsupported by the image. The problem: standard reinforcement learning rewards a response as a whole if the final answer is right, so those unsupported claims get credit alongside the correct answer. The team built a diagnostic that swaps in an altered image to check whether a model's claims are sensitive to what it's actually looking at, versus whether it just repeats the same claim regardless. They found some training methods, DAPO and VPPO, made models more responsive to images overall, but also made unsupported claims stickier - harder to retract even when wrong.
That's the more useful finding here: getting a model to look at an image more isn't the same as getting it to admit when the image contradicts what it said. The fix, called Persistence-Aware Credit Gating, docks reward specifically for claims that persist suspiciously often regardless of evidence, without needing anyone to label which claims are false. Applied to Qwen2.5-VL-7B, it nudged nine-benchmark average accuracy up by about 1-2 points across two training methods and also improved scores on HallusionBench, a benchmark built to catch exactly this kind of confident-but-wrong visual reasoning.
It's a narrow, technical fix, not a breakthrough - but it's a useful reminder that RL-trained models can learn to look busy without actually listening.