Giving an AI model time to reason before it answers does not make its decisions fairer. It just changes which biases show up.
Researchers tested three reasoning language models, QwQ-32B, DeepSeek-R1-Distill-Qwen-32B, and Qwen3-32B, on three widely used bias benchmarks: the Adult income dataset, the COMPAS recidivism tool, and a credit-scoring task. For each model, they compared quick non-thinking answers against answers produced after a chain-of-thought reasoning step, then checked for counterfactual flips, cases where changing a protected attribute like race or gender flips the decision even though nothing relevant to the outcome changed. Thinking did resolve some of those flips. But across all nine model-and-dataset combinations tested, it created roughly five new flips for every one it fixed, and did so with near-maximum model confidence.
That asymmetry is the interesting part. Reasoning models are often pitched as the more careful, more trustworthy option for exactly the kind of high-stakes calls used here, lending and criminal risk scoring. This study suggests reasoning can instead let bias compound: the researchers' own tracking tools show bias propagating and amplifying the deeper a model 'thinks,' rather than getting corrected out.
Showing its work, in other words, is not the same as getting the answer right.