A new AI model finds objects in images without reading them in order, and it barely loses accuracy for the speed.
A paper posted to arXiv describes GroundAnything, a 4-billion-parameter model built for visual grounding: the task of pinpointing exactly where an object or spatial relationship sits in an image. Most grounding models work autoregressively, predicting one token at a time in strict sequence, which is accurate but slow. GroundAnything instead uses diffusion, letting the model propose and refine many spatial guesses at once. The team trained it on public grounding datasets plus custom data pipelines, converted an autoregressive baseline into a diffusion model, then fine-tuned it with GRPO (Group Relative Policy Optimization), a reinforcement-learning method that ranks batches of the model's own outputs against each other instead of relying on a separate reward model.
Across 30 benchmarks, the model's autoregressive variant scored 72.42%, topping similarly sized rivals. The paper calls this "competitive with" a model it names GPT-6 Astra, which scored 71.35% - a name that reads like a mashup of OpenAI's GPT branding and Google's Astra project. The paper does not source or confirm what that model actually is, so that comparison should be read as unverified rather than a real head-to-head.
The more concrete result is the speed tradeoff: an optional self-speculative decoding mode claims a 4.51x speedup over standard autoregressive decoding for only a 0.74 percentage-point accuracy drop. That is the kind of math that matters for products like robotics or AR that need object locations in real time, not just a leaderboard score.
Whether the speedup holds outside benchmark conditions, and against competitors that are actually named, remains the open question.