Typing an angry version of your math homework into an AI chatbot can make it get the answer wrong, even if every number stays exactly the same.
Researchers built a benchmark called Temper-5400 to isolate whether emotional wording alone degrades quantitative reasoning in large language models. They took problems from three standard benchmarks: the grade-school math sets GSM8K and MultiArith, plus the science-reasoning benchmark ARC-Challenge, and rewrote each into an emotional version, injecting frustration, urgency, or enthusiasm while keeping every quantity and relationship untouched. The resulting 5,400 verified pairs were tested on eighteen models spanning 1-billion-parameter systems up to frontier-scale ones. Accuracy fell by 2 to 10 percentage points on the emotional versions compared with the neutral originals.
That is a real swing for a change that alters zero facts, just tone. The team also found that stripping the emotional language back out at inference time recovered most of the lost accuracy, while plain non-emotional paraphrasing caused no degradation at all, which points squarely at emotional framing, not surface rewording, as the cause.
Real users do not type in clean, emotionally neutral prose. A model that gets worse at arithmetic the moment you sound stressed is a usability problem hiding inside a tidy benchmark number.