A new paper lays out rules for picking useful synthetic training data instead of just generating more of it.
Researchers posted the work, called Training-Aware Target Coverage (TATC), to arXiv this week. They built a linear theory describing when synthetic data actually helps a model, how much of it to add, and what one more example is worth once a training set already exists. From that theory, TATC selects synthetic examples that fill gaps in a model's existing data, rather than adding more examples that look like what it already has. The team verified the theory on text and image data, then tested TATC on a real job: fine-tuning the small open model Qwen2.5-Math-1.5B-Instruct to solve GSM8K math word problems. TATC beat rival synthetic-data selection methods across every data budget tested.
This matters because labs have spent roughly two years leaning on synthetic data to make up for a shrinking supply of fresh human-written text, with uneven results: some report real gains, others quietly degrade on edge cases. A method for sorting good synthetic examples from noise, before spending compute training on them, is more useful than yet another synthetic dataset.
The catch: this is a 1.5-billion-parameter model on one math benchmark. Whether target-coverage selection holds up on messier tasks, or at the scale labs like OpenAI and Anthropic actually train at, is left untested.