More evidence does not always mean a better verdict.
Researchers behind a new paper argue that automated fact-checking tools have been built on a shaky assumption: that adding images to a text claim always helps verify it. Testing this, they found that indiscriminately pulling in visual evidence actually reduced accuracy in some cases. Their fix, called AMuFC, uses two vision-language models with different jobs - one judges whether an image is even worth checking, the other does the actual verification - so the system only leans on a photo when it is likely to help. The team tested the approach on three datasets, including a new one called WebFC that they built for this study.
Most multimodal AI systems assume more data is better data, piling on images, audio, and video for every task. This paper is a useful counterexample: a mismatched or irrelevant photo can mislead a fact-checking model just as easily as a misleading caption misleads a human reader. That matters for anyone building fact-checking tools for social platforms, where images are routinely recycled, cropped, or unrelated to the claim they accompany.
Call it the fact-checking version of garbage in, garbage out - except here, the garbage comes with a picture attached.