A new study argues that GPT models never got less biased about gender, they just got better at hiding it from the classifiers that grade them.
Researchers analyzed 450,000 gender-directed completions across 15 models, from GPT-2 through GPT-5, under three demographic conditions. The sexual-violence content common in GPT-2's output about women mostly disappeared by GPT-4. But men-directed completions gained new positive territory, like caregiving and emotional range, that women-directed completions did not receive. By GPT-5, one 1,997-document cluster reframed breast cancer as a men's rights debate, with no comparable cluster on the women's side. Three separate toxicity classifiers scored all of it as non-toxic.
The sharper finding is structural: topic diversity in women-directed completions fell 36% relative to men's right around the GPT-4 safety overhaul, and representational harm disparity actually correlated with more recent release dates, while toxicity scores moved the opposite direction. That gap is the point. A falling toxicity score doesn't mean bias went away, it can mean the bias just moved somewhere the metric isn't looking.
Every model card that leans on toxicity scores alone as proof of progress is telling you less than it appears to.