AI/ ai · llm-evaluation · benchmark · prompt-engineering

Angry Wording Alone Makes AI Reasoning Models Slip Up

A new benchmark finds emotional phrasing alone, with no numbers changed, cuts AI reasoning accuracy by 2 to 10 percentage points.

Typing an angry version of your math homework into an AI chatbot can make it get the answer wrong, even if every number stays exactly the same.

Researchers built a benchmark called Temper-5400 to isolate whether emotional wording alone degrades quantitative reasoning in large language models. They took problems from three standard benchmarks: the grade-school math sets GSM8K and MultiArith, plus the science-reasoning benchmark ARC-Challenge, and rewrote each into an emotional version, injecting frustration, urgency, or enthusiasm while keeping every quantity and relationship untouched. The resulting 5,400 verified pairs were tested on eighteen models spanning 1-billion-parameter systems up to frontier-scale ones. Accuracy fell by 2 to 10 percentage points on the emotional versions compared with the neutral originals.

That is a real swing for a change that alters zero facts, just tone. The team also found that stripping the emotional language back out at inference time recovered most of the lost accuracy, while plain non-emotional paraphrasing caused no degradation at all, which points squarely at emotional framing, not surface rewording, as the cause.

Real users do not type in clean, emotionally neutral prose. A model that gets worse at arithmetic the moment you sound stressed is a usability problem hiding inside a tidy benchmark number.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →