Audio-visual AI models sometimes talk themselves out of right answers the moment you show them a confusing picture.
Researchers describe a hallucination problem in audio-visual large language models: when audio alone is enough to answer a question correctly, adding video can still drag the answer down. In the paper's own example, a model correctly identifies a violin from audio alone, but its confidence drops once a video showing a guitar is added. The fix, called Relevant Evidence Decoding, is training-free. It runs a question-only pass to figure out whether audio, video, or their interaction actually matters for a given question, then uses pointwise mutual information to boost the contribution of whichever evidence type is relevant, rather than blending both modalities by default. Tested across three hallucination benchmarks and three AV-LLMs, it improved accuracy by up to 7 percent on CMM, 6.3 percent on AVHBench, and 3.8 percent on SVHalluc.
This matters because multimodal AI's selling point is combining senses, but the paper shows combining them can make a model dumber, not smarter, when one input is irrelevant or misleading. That is a basic reliability problem for anything built to watch and listen at once, like surveillance tools, video search, or accessibility software. A 1.5x hit to time-to-first-token is a real cost, but a model that gets quieter about correct answers when given noisy input is a harder sell than a slightly slower one.
Contrastive decoding already does something similar for single-modality vision-language models. This is that idea's first real stress test on two senses at once, and it suggests most AV-LLMs are currently bad at knowing which sense to trust.