Researchers have built a computational model that finds and removes a mark on its own face in a mirror — without being told to, and without a reward signal.
The model, described in a preprint, centers on what the authors call a "self-prior" — a Transformer trained on a simulated infant's normal multisensory experience: vision and body-position data, no touch. When a sticker appeared on the infant's face, the model flagged the anomaly as a mismatch against its learned body schema. That mismatch, processed through a framework called active inference, drove the model to reach up and remove the mark. In roughly 70 percent of trials, it succeeded. Expected free energy — a measure of surprise under the model — dropped after removal, which the authors take as evidence the self-prior was genuinely acting as an internal criterion for self versus non-self.
The mirror test has been the rough proxy for animal self-awareness since the 1970s, but it has always been criticized as behavioral: it records what an animal does, not why. This work proposes a single mechanistic answer rooted in the free energy principle, a theory from neuroscience that frames cognition as surprise minimization. If that framing holds, self-recognition is not a high-level cognitive achievement requiring special circuitry — it is a byproduct of any system that accurately models its own body.
The infant is simulated and the environment is controlled, so a 70 percent success rate is not a claim about machine sentience. What the paper is actually selling is parsimony: one mechanism, no reward shaping, emergent behavior. That is a testable hypothesis, and it will keep neuroscientists and AI researchers arguing productively for years.