AI/ ai · computer-vision · research · open-source

New Training Method Makes AI Trust Its Own Spatial Guesses

SpatialCORE, short for Spatially COnfident REasoning, rewards vision-language models for reasoning only from spatial guesses they are confident in.

A new training method teaches vision-language AI to stop guessing where things are and start pointing with confidence before answering questions about a scene.

The method, called SpatialCORE (short for Spatially COnfident REasoning), is a post-training framework for large vision-language models (LVLMs). These models already draw bounding boxes or masks around objects as part of their reasoning, but older training setups rewarded only the final answer being right, even when the box-drawing step was shaky or lucky. SpatialCORE changes the reward: it scores each predicted bounding box by how well it matches the object and how confident the model's own coordinate predictions are, then ties that grounding score to whether the final answer was correct. In testing, the approach beat other open-source and specialized spatial-reasoning models across several benchmarks and held up on data it had not seen during training, with code posted on GitHub.

This targets a specific failure mode: a model getting the right answer for the wrong reason, guessing correctly despite pointing at the wrong pixels. That gap between answer accuracy and grounding accuracy echoes a problem text-only language models have with chain-of-thought reasoning that sounds plausible but doesn't actually track the real steps. Making confidence itself part of the training signal is a more direct fix than hoping correct answers alone will weed out bad localization.

It's a research paper with a GitHub repo, not a shipped product - the real test is whether this holds up outside benchmark conditions, in the messy photos and diagrams people actually ask AI about.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →