Quality filters that decide what text is good enough for AI training have mostly worked in English. A new paper shows how to extend them to over 100 languages without starting from scratch.
The researchers took an existing English quality classifier, the kind of model that scores web text before it gets fed into large language model (LLM) pretraining, and built a multilingual version on top of it. They trained a small multi-layer perceptron, a lightweight neural network, on embeddings from a transformer encoder, using multilingual text as input and the English classifier's scores on machine-translated versions of that text as training labels. Tested on models at 1B, 3B and 8B parameters, the adapted classifier matched the performance of existing multilingual filtering baselines on standard benchmarks and did not hurt scores on tests of regional and cultural knowledge. It also generalized to languages it never saw during training, apparently picking up the English classifier's scoring logic through the translated labels alone.
This matters because quality filtering has been a quiet advantage for English and a few other high-resource languages, and a quiet disadvantage for everyone else. Low-resource languages rarely have enough annotated data to build their own filters, so their pretraining corpora tend to be noisier. Reusing an English classifier through machine translation, rather than demanding new labeled data for each language, could narrow that gap.
It is not a flashy fix, just a small network bolted onto existing embeddings, but pipeline plumbing like this often moves benchmark numbers more than the next architecture announcement does.