Picking which data trains a language model is itself becoming an AI problem.
Researchers are tackling a method called meta-learning for training-data selection, which trains a system to weigh examples by how well they serve a validation goal, instead of relying on hand-built heuristics. The obvious way to scale this up is to swap per-example weights for a full selection network. The authors found that doing so breaks down in practice: training grows unstable, and the network falls back on easy-to-learn shortcuts rather than genuinely useful examples. Their proposed fix, TESS, replaces the usual objective with something they call Pointwise Value Matching. In tests on safety-focused training and targeted instruction tuning, TESS carried its data judgments from small samples to full corpora, and from smaller models to larger ones.
Data selection is the unglamorous bottleneck behind every claim about bigger training runs: which examples you keep matters as much as how many you have. A selection method that fails to transfer means redoing expensive validation work every time the dataset or model size changes. That reusability, not raw data quality, is the specific problem TESS says it solves.
This is one arXiv paper, not a shipped tool. The real test is whether a lab outside the authors' group can reproduce the transfer claims on data TESS has never seen.