A new philosophy paper asks whether anything inside a large language model deserves to be called a mind, and its best candidate is a temporary one that exists only for the length of a conversation.
The paper, posted to arXiv's cs.AI section, tackles what it calls the individuation problem: which entities tied to a large language model, if any, should count as minds. To answer it, the authors turn to mechanistic interpretability, engaging with recent empirical work on persona vectors, persona space, and emergent misalignment, internal structures researchers have identified that correspond to distinct personas a model can adopt, including some linked to unsafe or broken behavior. They argue that attention mechanisms sustain a kind of continuity across a single conversation, which supports what they call the virtual instance view: a temporary, conversation-bound entity as the strongest candidate for an AI mind. They then introduce two further candidates, an instance-persona view and a model-persona view, built on that same body of persona research.
The paper does not settle on a single answer. It argues hardest for the virtual instance view but explicitly calls the two persona-based views promising alternatives, not runners-up to be waved off. That three-way hedge is itself notable: the question of whether anything in an LLM resembles a mind is now being fought with interpretability tools, not just intuition.
Before anyone argues a chatbot deserves rights, this paper suggests settling a smaller question first: what, if anything, is doing the talking from one exchange to the next.