A new paper gives a rigorous proof for something AI researchers have long suspected: language models are, at their core, compression engines.
The researchers show that large language models, modeled formally as next-token predictors, are mathematically equivalent to "monotone" compression algorithms - compressors that preserve the original order of the data they encode. The two can be converted into one another while losing almost no accuracy (an error gap of just 2). The team also pins down exactly when that equivalence requires monotonicity: if and only if a certain class of cryptographic one-way functions exists, the hard-to-reverse math problems that underpin modern encryption. A side effect of the proof strengthens an older, one-directional result from 2023 into a full two-way equivalence between a data distribution's unpredictability and its resistance to this kind of compression.
This matters because it explains, rather than just observes, a pattern other researchers already found: a 2024 ICLR paper showed LLMs work well as general-purpose compressors, and a 2024 COLM paper found that how well a model compresses text closely tracks how well it reasons and recalls knowledge. This new work supplies the missing theoretical bridge between those two findings, and ropes in cryptography as an unexpected bonus.
None of this makes any existing chatbot smarter overnight. It is a proof about the nature of next-token prediction, not a new training trick. But it hands researchers a cleaner vocabulary for a question that usually gets answered with vibes: why does squeezing text down well seem to go hand in hand with understanding it.