AI/ synthetic-data · llm-training · pre-training · arxiv

Study Finds Book Structure Boosts Synthetic Training Data

Researchers show organizing AI-generated text into full textbooks, not just rewriting passages, measurably improves model training results.

Turns out how you organize AI-generated training text matters as much as the text itself.

A new paper tests a common assumption in AI training: that synthetic textbook data helps mainly because of what it says, or how a passage is locally rewritten. Researchers built a pipeline that pulls source material from a pre-training corpus, groups it into topics, plans a table of contents, and assembles full books from it, generating 686,000 textbooks totaling 32 billion tokens across more than 15,000 subjects. Swapping these books into a mid-training mix in place of natural books improved downstream performance by 1.09 points on average. To isolate why, they ran controlled comparisons: splitting the same generated text into standalone sections lost most of the gain, randomly concatenating unrelated sections also underperformed, and simply rewriting individual documents without any book-level planning fell further behind.

This matters because most synthetic-data work chases better content or better per-passage style, assuming organization is just packaging. This study suggests the opposite: deliberately structuring related material into a coherent, book-like whole is doing real work, not just making data look tidier. That is a cheap lever anyone building training pipelines could pull, independent of which generation model they use.

The result also held on Llama3-8B, not just whatever model the researchers used internally, which is a decent sign this is not a one-model fluke. Worth remembering, though: this is one architecture family and one mid-training setup, not proof the effect holds everywhere.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →