Ask an AI model to screen a tenant, and the dialect used in the application might matter as much as its content.
Researchers ran a matched-guise experiment, a classic sociolinguistics method, on ten open-weight large language models. They built 260 meaning-matched sentence quadruples, the same message rewritten in Standard American English, African American Vernacular English, Nigerian Standard English, and Nigerian Pidgin, then measured which housing-related adjectives each model rated as more probable across three scenarios: tenant screening, neighbor acceptance, and roommate selection. Across all ten models, African American Vernacular English and Nigerian Pidgin drew more negative adjectives than Standard American English, with Nigerian Pidgin penalized hardest. Nigerian Standard English behaved differently, scoring better than Standard American English in formal tenant screening but losing that advantage as scenarios grew more socially intimate, such as picking a roommate.
These biases live in the models' internal probability distributions, not in the text they generate, so they survive the alignment work that scrubs slurs and obvious stereotypes from outputs. Housing is where that abstraction turns concrete: if any of these ten models feeds a screening tool, the bias could echo the same dialect-based housing discrimination already documented among human landlords, just laundered through a probability score instead of a gut reaction. The study also found each dialect is penalized through its own distinct stereotype cluster rather than one generic penalty for sounding nonstandard, meaning a single bias fix will not cover all four varieties.
It is a useful check on bias research that stops at African American Vernacular English. Nigerian Pidgin, a variety almost absent from prior covert-bias studies, came out worse than all three other dialects tested, including the one everyone already knew to worry about.