AI/ ai · benchmarks · computer-vision · multimodal-ai

New Benchmark Shows AI Models Fail Basic Visual Logic Tests

Hob-VL tests whether AI can combine simple visual facts with Boolean logic, and most models score barely better than a coin flip.

AI models that can recognize a cat in a photo still struggle to reason logically about what they are seeing.

A new benchmark called Hob-VL tests whether vision-language models can combine simple visual observations using Boolean logic - AND, OR, NOT, and nested combinations - rather than just spotting objects. The dataset includes 6,000 human-verified Yes/No questions built from combinations of ten visual statements, spanning 1,000 generated scenes and 46 labeled photographs, plus 1,000 questions asking models to identify the one object matching a given description. Questions are deliberately built with misleading local cues and nested logical operations, presented in both symbolic and plain-language form. Across eight model setups with little or no step-by-step reasoning enabled, Boolean-question accuracy landed between 48.52% and 50.57% - essentially a coin flip on a yes/no task - while object-identification accuracy topped out at 43.0%. Even a reasoning-enabled GLM configuration only improved unevenly, with researchers reporting persistent errors and inconsistent answers to logically equivalent questions.

That gap matters because compositional reasoning, not just object recognition, is what separates a model that can describe a photo from one that can be trusted to act on what it sees - in search, robotics, or any tool that chains visual facts into a decision. Near-random scores on a benchmark built from simple logical combinations suggest current vision-language models are still pattern-matching their way through images rather than reasoning about them.

A model that answers 'is there a red car or a blue bike' correctly but flips to wrong on the logically identical 'is it false that there is neither a red car nor a blue bike' was never really reasoning - it was guessing with good vocabulary.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →