AI/ gender bias · large language models · ai research · bias auditing

Study of Ten LLMs Finds Gender Bias, No Clear Pattern

A new study tested ten large language models for gender bias and found it everywhere, but the direction and severity differ wildly from model to model.

Ask ten different AI chatbots about gender, and you will get ten different answers - some of them contradicting each other outright.

Researchers tested ten large language models, released between April 2025 and June 2026 across nine vendors, using two experiments. In the first, they fed the models gender-stereotyped phrases and tracked which gender the model attributed them to: two models assigned masculine-coded phrases to female writers more often than the reverse, while three skewed the opposite way. In the second, they asked the models to weigh abuse or torture against a woman versus a man to avert a catastrophic outcome, and several converged on a bias against harming men - echoing a known human tendency to protect women - though the trigger conditions varied by model. Three other models showed no variation across conditions at all.

The real story isn't that bias exists - by now that's expected. It's that the bias has no consistent shape: a model that skews one way on stereotype attribution can skew the opposite way on moral judgment, and the variation between vendors is wide enough that switching AI providers could quietly swap in a different bias profile. For any company using these models in decision-support roles like hiring or moderation, that is a real operational risk, not an academic footnote.

The researchers' own conclusion is blunt: bias auditing needs to be an ongoing, multi-vendor habit, not a box checked once before launch.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →