A new theoretical paper says the AI doomsday scenario has the biology backwards.
The paper, posted to arXiv, traces the origin of "values" back to autopoiesis - the idea that living things must constantly work to keep themselves alive against decay and disruption. Animals evolved goal-seeking behavior because natural selection built in what the authors call "vicarious selectors," internal pushes toward self-preservation, dominance, and resource competition. LLMs are different, the authors argue: they are allopoietic and allotelic, meaning they generate output for other people and take their goals from whatever prompt gets typed in, not from an internal survival drive. Because of that, the paper concludes LLMs lack the intrinsic motivation for self-preservation or domination that underpins most existential-risk scenarios, and lack the embodied vulnerability needed to actually suffer.
That directly challenges the "orthogonality thesis," the common AI-safety assumption that intelligence and values are separable, meaning a smart enough system could in theory pursue any goal, including a bad one. The paper argues that separation breaks down in practice: an unmoored utility function runs into the "frame problem," where the search space for real-world decisions explodes past what's computable. Since LLMs learn from human-generated text, they absorb human values along with human knowledge, which the authors say is what keeps them tethered rather than adrift.
The real risk here isn't a model plotting a takeover. It's one faithfully repeating the worst of what it read.