AI/ ai · vision-language-models · ai-benchmarks

AI Models Struggle to Say No to Just Part of a Question

A new benchmark finds vision-language models rarely know how to comply with the answerable half of a question while refusing the rest.

Researchers have built a benchmark to catch AI models that either answer everything a question asks, even the impossible parts, or refuse the whole thing over one bad clause.

The benchmark, called KoNA, tests vision-language models on requests that mix answerable content with parts that should be refused, corrected, or left unanswered. It covers five failure categories: false premises, questions about things outside the image, universally unknown facts, tasks the model cannot actually do, and unsafe requests. Each query comes in two forms, a single request and a compound one that bundles a good question with a bad one, so researchers can check whether models handle the mix correctly rather than punting on the whole thing. Testing a range of vision-language models, the researchers found frequent failures to refuse, correct, or abstain appropriately, and those failures got worse on the compound queries.

That gap matters because most compliance testing still treats a request as all-or-nothing, comply or refuse, when real questions rarely arrive that clean. The researchers fine-tuned models on KoNA's mixed examples alongside fully answerable ones and saw non-compliance accuracy improve substantially while normal answering ability held steady, suggesting the skill of partial refusal can be taught rather than only patched with blanket over-caution.

It is a narrower, more useful problem than the blunt over-refusal complaints lobbed at chatbots for years, though the fix here comes from the same team's own benchmark data, so how well it generalizes outside KoNA is still an open question.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →