AI/ ai-benchmarks · vision-language-models · llm-reasoning · active-learning

Benchmark Tests Whether LLMs Know When Vision Models Are Wrong

VisualNoiseQA forces text-only language models to query an unreliable vision model and decide for themselves when to trust, re-ask, or give up.

A new benchmark checks whether AI reasoning models know when to doubt their own eyes.

VisualNoiseQA, described in a new arXiv paper, makes a text-only LLM solve visual question-answering problems without ever seeing an image directly. Instead, it repeatedly queries an off-the-shelf vision-language model (VLM) treated as a noisy sensor, pulling multiple samples per query and getting an uncertainty signal from how much those samples agree with each other. The LLM has to decide what to ask next and when it has enough evidence to stop. The researchers built the 1,000-question set automatically, pulling from existing VQA sources and keeping only cases where two different noisy VLMs gave inconsistent answers that a human could still figure out, spanning basic perception, chart reading, and knowledge-heavy questions.

Most active-reasoning benchmarks quietly assume that whatever a tool or sensor reports back is true, which is not how deployed systems behave - cameras misfire, OCR garbles text, and classifiers hedge. By exposing a calibrated uncertainty signal instead of hiding the noise or ignoring it, VisualNoiseQA isolates a specific, underexamined skill: knowing when to re-check a source versus when to commit to an answer. That is closer to how an autonomous agent using real cameras or flaky APIs would actually have to operate.

The abstract stops short of naming which model handled the uncertainty best, so the interesting part - whether current LLMs are any good at not trusting a shaky witness - is still an open question.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →