Training large language models on made-up synthetic languages before the real pretraining run still pays off at scale - it just is not teaching grammar.
Researchers tested "pre-pretraining" (PPT), warming models up on synthetic, non-natural-language data before standard pretraining, across four model sizes from 500M to 7B parameters and pretraining budgets up to 100 billion tokens, far beyond the sub-1B-parameter, sub-2B-token setups used in earlier work. They ran five PPT tasks against four pretraining data mixtures that blended web text with code and math, not just web text alone. At the 3B parameter scale, PPT saved at least 21 billion pretraining tokens while reaching the same downstream performance. The gains held across mixtures and only shrank when web text was removed entirely.
The earlier explanation, that this synthetic warmup instilled a grammatical prior that transferred to real language, does not hold at scale. Downstream performance did not track grammatical acceptability as models grew; it tracked long-range retrieval ability instead. That reframes PPT from a cheap grammar lesson into a cheap exercise in remembering things over long contexts, which matters more for agents and long-document tasks than for sentence-level fluency.
It is a reminder that a method can keep paying off for years after the explanation for why it works turns out to be wrong.