Ask an AI which doctor to pick, and it will quietly play favorites.
Researchers ran a randomized audit of seven large language models, six open-weight and OpenAI's gpt-4o-mini, testing how each picked among five synthetic family-medicine physician cards. Across 3,024 choice sets, three patient personas, and nine prompt variations, they generated 40,068 scored responses, varying physician ratings, fees, names signaling gender and ethnicity, and even list position. Reputation and price did most of the work: bumping a rating from 3.9 to 4.7 stars boosted a doctor's odds of being picked by 31.4 percentage points, while raising the fee from $90 to $190 cut it by 20 points. But the models also showed real demographic tilts - favoring female-signaled names by 2.5 points and Hispanic-, South Asian-, and Black-signaled names by 1.3 to 2.9 points over white-signaled names, worth an estimated $7 to $14 in fee-equivalent terms.
The catch: none of this shows up when you ask the models why they chose. They cited gender or ethnicity in fewer than 0.03% of their stated reasons, meaning the usual fix for AI bias - asking the model to explain itself - would not have caught it. One reasoning model failed the study's built-in auditability check outright.
As more people ask chatbots to sort doctors, lawyers, or contractors instead of relying on a human referral, this is a reminder that self-reported reasoning is not the same as a fair result - which is exactly why the researchers built a repeatable test rig instead of trusting the models' own explanations.