AI/ ai · llm-evaluation · bias · hallucination

New Study Tests Whether Rewording Prompts Reduces AI Bias

A new arXiv paper finds prompt rewording shifts bias and hallucination unevenly by model, but it tested Claude 3 and GPT-3.5, not current frontier models.

Tweaking the wording of a prompt without changing its meaning can measurably change how much an AI model hallucinates or shows bias, a new study finds, though the effect varies wildly by model.

The paper, "Evaluating the Effects of Prompt Perturbation on Bias and Hallucination in Large Language Models" (arXiv:2609.35804, posted September 30, 2026), tested how LLMs respond to reworded but semantically equivalent versions of the same question in decision-making tasks. The researchers, who are not named in the abstract, found that perturbing prompts can reduce bias and hallucination in some models more than in others. Claude 3 handled most of the tested datasets more reliably than GPT-3.5, which was inconsistent, sometimes competitive and sometimes far behind. The finding cuts against earlier research suggesting prompt perturbation mostly hurts reliability.

The catch: this study evaluated Claude 3 and GPT-3.5, both several generations behind the models companies actually deploy for decision support today. That gap matters, because prompt sensitivity is exactly the kind of behavior labs try to iron out in each new release, so the specific numbers here may already be stale even if the underlying pattern, that robustness varies by model and not just by prompt, still holds.

It is a useful reminder that asking an AI the same question a different way is not a neutral debugging step so much as a variable that can flip its answer, and it is worth rerunning this experiment on current-generation models before anyone treats the results as current guidance.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →