A new research paper makes an unusual case for AI vision: models learn 3D scene structure better when their decoders are made weaker, not stronger.
The researchers built SNAP, a self-supervised encoder-decoder transformer for novel view synthesis - the task of generating images of a scene from new camera angles. They argue that existing methods produce poor geometric representations not because they lack training signal, but because of two subtle architectural habits: decoders expressive enough to paint in missing details on their own, and training targets based on raw pixels rather than learned features. SNAP addresses both by using a pose-conditioned local decoder with deliberately limited expressive power, and by reconstructing targets in latent space instead of pixel space. The resulting model is task-agnostic, meaning it isn't built or tuned for one specific downstream job.
That generality is the point. Across five separate tasks - visual localization, pose estimation, point correspondence, depth estimation, and robot manipulation - SNAP holds its own against both specialized geometry-supervised systems and other self-supervised approaches, despite using less compute and data. Its patch-level features also show viewpoint invariance that approaches heavily supervised models, and they degrade more gracefully than standard 2D representations when the camera shifts.
It's a tidy counterexample to the usual scaling instinct in AI research: sometimes the fix for a mediocre representation is a smaller, dumber decoder, not a bigger one.