AI/ ai · computer-vision · visual-grounding · research

New Framework Lets a 9B Model Match 241B Rivals at Grounding

CoEvolve, a 9B visual-grounding model, matches accuracy from models up to 241B parameters by making its reasoning steps explicit and editable.

A 9-billion-parameter model now matches object-finding accuracy from AI systems up to 241 billion parameters in size.

Researchers introduced CoEvolve, a visual grounding model that finds objects described in plain language by drawing a bounding box around them. Instead of guessing the box in one shot, it splits the task into two stages: Region-Evolution Reinforcement narrows candidate regions step by step, and a separate module, Bidirectional Denoising Refiner, cleans up the resulting coordinates while leaving the reasoning text untouched. The team tested the 9B model on ordinary photos and satellite imagery, where it matched grounding accuracy from models as large as 241 billion parameters. In a separate experiment, the researchers deliberately corrupted inputs to simulate bad initial guesses, then ran a single refinement pass; that test recovered more than 27 percentage points of box-overlap accuracy, a different measurement from the parameter-count comparison above.

This matters because grounding errors are usually invisible until the final answer is wrong, with no way to tell whether a model mislabeled the region or just drew a sloppy box. Making each reasoning step produce an explicit, editable coordinate state means errors can be caught and fixed mid-process, which matters for robotics and augmented-reality systems that need reliable real-time localization, not just a nice caption. A model this size holding its own against systems with 241 billion parameters also hints that smarter architecture, not just scale, can close performance gaps in vision tasks.

That said, the comparison comes from the authors' own preprint, which has not been peer reviewed, and a 9-billion-parameter model is still too large to run on a phone, so any efficiency win here is relative, not absolute.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →