AI/ ai · clinical ai · ai fairness · research

Fairness Audits of Medical AI Need a Noise Floor First

A new benchmark shows clinical AI agents flip decisions on their own, so fairness audits must subtract that baseline noise before blaming bias.

A new audit tool shows some of the bias flagged in clinical AI fairness tests might just be randomness.

Researchers tested clinical AI agents for counterfactual fairness, the standard check of whether a recommendation changes when only a patient's demographic label changes. But first they reran the same scenario ten times with nothing altered, across sixteen synthetic patient vignettes, and found the agent's own action still flipped 8.7 percent of the time. That noise floor ranged from 2.2 percent for intensive-care escalation calls to 17.9 percent for controlled-substance caution calls, a category the researchers note has no clear operational criteria to begin with. Across six models from five vendors, the floor spanned 2.5 to 23.7 percent, with no consistent pattern tied to model size, vendor, or hosting setup.

The math explains why this matters. For a yes-or-no decision, the flip rate expected from pure noise equals that floor, and a real demographic bias only adds its square on top. That means a flip rate sitting inside the noise floor is not evidence of fairness or bias; it is just the model disagreeing with itself. The floor is not fixed, either: majority voting over five draws cut it by 39 percent, and setting temperature to zero eliminated disagreement for three of four self-hosted models but not for a hosted one.

The team is releasing the harness, called FairMedAgent, along with its protocol, vignettes, and analysis scripts, so other teams can measure their own agent's floor before claiming the flip rate they see is bias.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →