When you ask a language model to rewrite a scientific or medical text, it tends to make the claims sound more certain than they actually were.
A new study introduces a metric for "certainty distortion," meaningful shifts in how confidently a claim is expressed during transformations meant to preserve meaning, like paraphrasing or summarizing. Testing multiple model families and sizes on scientific and medical communication tasks, the researchers found certainty distortion in up to 75% of outputs, and models were 1.5 to 2 times more likely to inflate confidence than to soften it. The effect compounds with repeated editing: in the medical domain, claude-haiku-4.5 increased certainty in 20% of examples after a single rewrite, rising to 40% after five iterations of paraphrasing. Prompt-based instructions to preserve hedging reduced the distortion but did not eliminate it.
This is not the usual AI failure mode people worry about. It is not a hallucinated fact, it is a shift in tone that nobody would catch without lining the rewrite up next to the original. In medicine, "may indicate" quietly becoming "indicates" is exactly the kind of change that could nudge a treatment decision or a public health message, with no factual error anyone could point to.
The paper's framing captures it well: the failure here is not the model getting facts wrong, it is the model getting confident wrong, and confidence is often the only thing a reader has time to scan for.