AI/ ai · benchmarks · multimodal-ai · spatial-reasoning

New Benchmark Shows Best AI Model Solving Only Half of 3D Puzzles

PolyComp's polycube puzzles cap the top AI model at 50 percent accuracy, with a rival scoring barely above random guessing, at real per-problem cost.

A new benchmark shows the best multimodal AI models can barely combine two cube-shaped puzzle pieces correctly half the time, and it costs real money to watch them try.

Researchers built PolyComp, a procedurally generated and independently verified benchmark of 120 problems spanning four geometry families. Each problem asks a model to pick, from four options, the one pair of polycube pieces that combines into a target 3D solid, shown in three different image formats for a total of 360 presented problems per model. Random guessing scores 25%. Running each model at its highest effort setting, GPT-5.6 Sol topped the field with 50.0% accuracy at $0.951 per problem, Claude Fable 5 followed at 39.4% for $0.701, and Gemini 3.1 Pro Preview landed at 27.5% for $0.350, barely above chance.

Spatial reasoning like this underpins real work: robotics, CAD tools, AR overlays, anything that needs a machine to understand how physical parts fit together. The benchmark also found accuracy swings more with the shape of the puzzle than with how the images are presented, meaning the bottleneck is genuine geometric understanding, not input formatting. Price did not track with performance either: the cheapest model scored worst, but that is no compliment to the pricier ones, since a coin flip beats two of the three.

Fifty percent right, at a dollar a guess, is not the number any of these labs would want headlining their own announcement.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →