Ask an AI the same question twice, swap only the advertiser's name, and you can get a different answer.
Researchers ran a systematic fairness audit of large language models used to judge whether an ad matches a search query. They tested GPT-4o as a general-purpose relevance judge and a Qwen-7B model fine-tuned specifically for relevance prediction, feeding both real query-ad pairs pulled from advertising logs plus synthetic queries built to probe employment, housing, and credit scenarios. Simply changing the advertiser's identity or the input language shifted relevance scores for both models. Synthetic demographic tests also turned up stereotype-shaped patterns, especially around gender and occupation. The team tried fixing this at inference time and during training, with mixed results tied to how advertiser labels show up in training data.
This matters because ad relevance models increasingly decide which businesses get shown to which searchers, and that's exactly the kind of decision regulators scrutinize when it touches jobs, housing, or credit. An AI that quietly favors certain advertiser names or stumbles on gendered job titles isn't a rounding error -- it's a mechanism for reproducing old-fashioned discrimination at algorithmic speed.
It's the ad-tech version of a problem AI researchers have been chasing in hiring and lending tools for years: bolt a general-purpose language model onto a high-stakes ranking decision, and its training-data baggage comes along for the ride. The paper's finding that mitigation only works when advertiser identity isn't actually relevant to the query suggests there's no easy patch here, just an ongoing tuning problem.