A new benchmark finds that closed-source AI models are more likely to produce biased answers to English speakers than to speakers of Hindi, Bengali, Tamil, Marathi, or Hinglish.
Researchers built INCLUDE, a 2,604-prompt benchmark testing for Indian socio-cultural bias across six languages: English, Hindi, Bengali, Marathi, Tamil, and Hinglish, the common Hindi-English code-mix heard across Indian cities. They ran ten open-source and closed-source LLMs through it, producing 14,988 bias scores. Bengali came out worst for open-source models. English came out best for open-source models, but worst for closed-source ones - a reversal the researchers flag as a key finding.
That reversal matters because safety alignment for most large models is built and tested almost entirely in English, on the assumption that English is the safest default. For closed-source models, the kind deployed in voice assistants and dialogue systems across India's linguistically diverse population, that assumption breaks down exactly where it is supposed to hold. Users speaking English to these systems got less-safe outputs than the safety-training story would predict, while other Indian languages fared better on the same closed models.
It is a reminder that aligned in English and aligned are not the same claim, and vendors marketing global safety on the strength of English-language red-teaming should expect that pitch to get tested in exactly the multilingual, code-mixed conversations they did not train for.