You can predict how well training an AI model in one language will help it in another without running a single model.
Researchers built a random forest model using only typological features, grammar traits, word order, and other properties cataloged in existing linguistic databases, to predict cross-lingual transfer scores. They tested it against a 24-language transfer matrix that prior research had to generate through expensive multilingual pretraining runs. The typology-only model scored a 0.705 correlation and an R-squared of 0.49 on leave-one-language-out tests, beating a non-typological baseline that managed 0.62. The result held up even when researchers withheld entire language families or scripts, which rules out the model simply memorizing those groupings.
Teams building models for lower-resource languages usually pick a high-resource source language by running costly training experiments or by defaulting to whichever language has the most available data. This method offers a free screening step that uses traits already logged in public databases, letting teams narrow candidates before spending compute on any actual training runs. The researchers also found that typology itself is not skewed by the resource-and-script bias that otherwise distorts which source language looks best, unlike the current data-driven approach.
A correlation of 0.7 is a useful shortcut, not a verdict, so anyone shipping a language model would be wise to treat this as a first filter rather than the final word.