A new paper gives a precise, testable answer to a question the AI field mostly guesses at: which downstream tasks actually survive self-supervised pretraining.
The researchers studied same-instance self-supervised learning, the method behind systems like SimCLR and many vision and audio encoders, which trains a model by making it agree with itself across two slightly different views of the same input. They define "semantic recoverability" as how much of a task's signal is captured in the learned representation, then show it has an exact mathematical relationship to known measures like class-distance-normalized variance and few-shot nearest-centroid classification accuracy. For a standard two-view training objective, they prove the model's optimal representation spans a specific set of "spectral" directions shared across views, meaning a task is preserved only if its signal happens to live in that subspace. They back this up with experiments on synthetic and real datasets across several SSL methods.
That matters because SSL pretraining is usually treated as a black box: train on unlabeled data, then hope the resulting features transfer to whatever task you care about. This work suggests transfer success isn't random or purely a function of data scale - it's determined by whether a task's structure aligns with a predictable spectral subspace, which could let engineers estimate transferability before spending compute on downstream fine-tuning.
It's a theory paper, not a tool - don't expect a "recoverability score" in your ML framework's documentation next week. But it's the kind of math that eventually turns pretraining strategy from folklore into something you can actually plan around.