A new AI training method learns the 3D shape of a scene without ever trying to redraw it.
Researchers describe Poincar3, a self-supervised method that learns from multiple camera views using self-distillation instead of RGB pixel reconstruction. It pairs masked patch and image-level distillation with a teacher model that sees extra viewpoints, so the system trains from scratch with no explicit 3D labels. In tests on correspondence estimation, camera pose estimation, and 3D reconstruction, it beat prior single and multi-view self-supervised models, including DINOv3, MuM, and Muskie. A lightweight add-on called a Poincare adapter also let the model estimate camera motion more precisely than existing self-supervised approaches.
Most vision models still learn by trying to rebuild pixels, which ties geometry to appearance - a model that memorizes how a cat looks isn't necessarily learning how space works. By skipping pixel reconstruction and distilling from multiple viewpoints instead, Poincar3 points toward representations that separate what something looks like from where it is, which matters for robotics, AR, and any system that needs to navigate a space rather than just caption it.
Henri Poincare argued over a century ago that a static observer can't develop a sense of space; this paper's answer is to stop trying to get there through pixels at all.