AI/ synthetic-data · llm-fine-tuning · machine-learning · ai-research

A Smarter Way to Pick Synthetic Training Data for LLMs

A new theory-backed method called TATC picks synthetic training examples that actually help a model, and beat rivals fine-tuning a math model on GSM8K.

A new paper lays out rules for picking useful synthetic training data instead of just generating more of it.

Researchers posted the work, called Training-Aware Target Coverage (TATC), to arXiv this week. They built a linear theory describing when synthetic data actually helps a model, how much of it to add, and what one more example is worth once a training set already exists. From that theory, TATC selects synthetic examples that fill gaps in a model's existing data, rather than adding more examples that look like what it already has. The team verified the theory on text and image data, then tested TATC on a real job: fine-tuning the small open model Qwen2.5-Math-1.5B-Instruct to solve GSM8K math word problems. TATC beat rival synthetic-data selection methods across every data budget tested.

This matters because labs have spent roughly two years leaning on synthetic data to make up for a shrinking supply of fresh human-written text, with uneven results: some report real gains, others quietly degrade on edge cases. A method for sorting good synthetic examples from noise, before spending compute training on them, is more useful than yet another synthetic dataset.

The catch: this is a 1.5-billion-parameter model on one math benchmark. Whether target-coverage selection holds up on messier tasks, or at the scale labs like OpenAI and Anthropic actually train at, is left untested.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →