A new computer-vision model called GroundingPI claims state-of-the-art results at the unglamorous but critical task of telling an AI exactly where an object is.
GroundingPI is a 4-billion-parameter model trained to output precise points and boxes for objects in a scene, including small or cluttered ones, quickly enough for closed-loop robot control. Researchers trained it with multimodal pretraining, supervised fine-tuning, and reinforcement learning, then benchmarked it against 44 other models across 34 tests spanning 11 perceptual skills, where it averaged 73.68% accuracy versus 71.54% for the next-best, larger general-purpose model. Used as the vision backbone for robot manipulation, it beat every rival system tested on RoboTwin 2.0's hardest out-of-distribution scenarios, by up to 24.8%, and on RoboCasa-GR1 it outperformed baselines trained on 75% of demonstration data despite being trained on just 50% itself. Plugged into a self-driving perception stack on the nuScenes dataset, it posted an average open-loop position error of 0.296 meters.
Most robot and driving AI systems borrow their vision from general-purpose vision-language models that were never built for the fiddly task of locating one specific object in clutter fast enough to act on it. GroundingPI's results argue that a smaller, dedicated grounding model can beat scaled-up generalists at exactly the job physical AI needs most, and that doing more with less training data is possible when the model is built for the task rather than adapted to it.
These are the authors' own benchmark numbers, not independent verification, and the rival models used for comparison are not identified by name or maker, so treat the margins as a starting point for scrutiny rather than a settled ranking.