AI/ ai research · computer vision · 3d reasoning · vision-language models

Researchers Find a Fix for AI's Shaky Sense of 3D Space

A new technique called GeoLatent untangles how AI models represent position and shape in 3D, pushing accuracy past prior benchmarks on spatial reasoning tests.

A new AI architecture called GeoLatent gets measurably better at figuring out 3D space from a flat photo, by refusing to cram position, direction, and overall shape into one blurry representation.

Researchers built GeoLatent by combining a method called Common-Residual Geometry Alignment (CR-GEO) with a training scheme they call routed optimization, aimed at a known weak spot in vision-language models: reasoning about 3D layout from 2D images. Earlier "decomposed spatial latent" approaches already split geometry signals into position, direction, and global shape, but the shape representation kept collapsing toward one dominant direction instead of staying varied. CR-GEO fixes that by separating shared geometric signal from residual, teacher-supplied detail. Routed optimization then trains geometry and language together, temporarily forces answers through that geometry bottleneck, and restores full image access once the representation holds - after which, per the paper, the geometry's "effective rank," a measure of how varied its internal representation is, rose from 1.00 to 3.87.

The payoff shows up in testing: GeoLatent scored 73.0% on SPAR-Bench and 72.1% on SPBench, ahead of previously reported methods on both. The more telling number is what happens when the shortcut is removed - blocking the model from reading its own geometry latents mid-task dropped direction accuracy from 89.1% to 25.8%, which suggests the representation is doing real work, not padding a benchmark line.

Spatial reasoning is the gap that keeps AI assistants from reliably telling you what's behind a shelf or whether a parking spot will fit a car; this is incremental, benchmark-bound progress toward closing it, not a solved problem.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →