Fine-tuning a small open-source language model made it less likely to change its emergency-room triage call based on a patient's race, income, or insurance status than the bigger, fancier version it was built on.
Researchers audited ten open-source LLMs on pediatric Emergency Severity Index predictions, including Qwen2.5-7B, Qwen2.5-14B-Instruct, a QLoRA fine-tuned Qwen2.5-7B, MedGemma variants, MedLLaMA2-7B, and GPT-OSS-20B and 120B. They fed each model near-identical clinical vignettes that differed in only one injected detail, things like a child's socioeconomic background, housing status, or how they arrived at the hospital, and tracked how often the triage level changed anyway. The fine-tuned Qwen2.5-7B shifted its answer in just 5.27% of these counterfactual pairs, versus 16.02% for the base Qwen2.5-7B it was tuned from, with a correspondingly smaller mean absolute shift (0.0534 versus 0.1706). Several larger models and medical-domain-pretrained models, the ones you'd expect to do better, showed bigger shifts instead.
That non-relationship between size, medical training, and fairness is the real finding here. It suggests hospitals treating "the biggest model" or "the one trained on medical data" as a built-in bias safeguard are checking the wrong box; sensitivity to irrelevant patient details has to be measured directly, not assumed from a spec sheet. The study's stratified analysis also found consistent directional patterns in who gets over- or under-triaged, not just random noise.
Cheap, lightweight counterfactual audits like this one could become a standard pre-deployment checklist item, much like bias testing became routine for hiring algorithms, except a missed bias here gets measured in ESI levels, not rejected resumes.