AI/ ai · llm-training · benchmarks · research

Correct Answers Aren't Always the Best AI Training Data

A new framework shows that AI training data which looks most correct is not always the data that most improves a model.

Picking AI training data by how correct it looks turns out to be the wrong instinct, according to new research.

A paper called FAER tackles a problem in how language models get further trained after their initial build: when old model outputs are cached and replayed for more training, the usual selection methods rank them by format compliance, confidence scores, or freshness, not by whether they actually help the model learn. Researchers tested this on GSM8K math problems using Qwen2.5-1.5B-Instruct, a 1.5-billion-parameter model, over 128 training updates. A selector built around format feedback picked responses that were correct 69.53% of the time, far more often than the 35.94% correctness rate of FAER's own baseline selector. Despite picking worse-looking data, FAER's baseline still trained a better model, scoring 0.6329 on the paper's quality measure versus 0.6037 for the format-feedback selector and 0.5482 for random selection.

That gap matters because most post-training pipelines assume correct-looking outputs make better training examples. FAER's results say that assumption does not reliably hold; what makes a training example useful is a different property than how accurate its answer is. A metadata-only version of FAER's calibrated selector pushed quality to 0.6476 on average across eight test runs, and a fuller version of the same approach reached 0.6624.

None of it is free. The metadata-only version needed 189,642 tokens of total compute to produce that result, even though only 63,276 of those tokens went toward the actual target task; the rest quietly covered tuning the selector itself, a cost easy to leave out of a headline number.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →