AI/ ai bias · reproductive health · llm evaluation · research

AI chatbots predict more judgment fear for Black abortion patients

A new evaluation method found four of five AI models scored Black personas higher on fear of judgment after abortion, reversing a real-world stigma pattern.

A new evaluation method found a racial skew in how AI chatbots predict judgment around abortion, even as their actual advice stayed the same for everyone.

Researchers introduced "behavioral coherence evaluation," a design-time method that checks whether an AI model's answers on a sensitive topic hold together the way a validated psychological instrument predicts they should. They used the Individual Level Abortion Stigma Scale to prompt five large language models to complete stigma questionnaires as 627 different personas, then had five reproductive-health experts review the flagged inconsistencies. The models consistently scored personas lower on self-judgment than on worries about how others would judge them, and for most models, worry-about-judgment became the single highest-scoring stigma dimension - even though that same dimension scored lowest in the human reference sample the scale was built on. Four of the five models reversed that reference pattern specifically for Black personas, generating significantly higher worry-about-judgment scores than the reference data would predict.

That racial gap didn't show up in the advice itself. Every model defaulted to recommending extreme secrecy after an abortion regardless of persona, even as stigma patterns varied across personas - a one-size-fits-all script that expert reviewers said misses what actually matters: relationship safety, legal risk, and access to trusted support.

It's a reminder that a chatbot can sound even-handed on a sensitive topic and still bake in a racial disparity that only shows up once you test it against an instrument built to measure exactly that.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →