Ask a chatbot what is not the capital of Spain, and it might still say Madrid.
A new benchmark testing both open-source and closed-source LLMs on negated questions found that in 37 to 71 percent of cases, the models answered with the exact word the question told them to exclude. Researchers didn't stop at the error rate; they opened up the models to see what was happening underneath. They found specialized attention heads and MLP neurons that try to handle negation by suppressing the original answer and boosting a different candidate from the same category, so 'Madrid' gets pushed down while 'Paris' gets pushed up. That's the opposite of how people seem to process negation, where the original answer is used as information to figure out what to rule out, rather than just being erased.
The mismatch explains the failures: when suppression is too weak, or the model already favors one answer, it falls back to the very thing the question ruled out. The researchers used that mechanistic finding to build a training method that forces bigger preference shifts for answers the model was most confident about, and it cut negation errors with less damage to the model's other skills than standard fine-tuning.
Negation is a small piece of language, but it's the kind of small piece that breaks trust fast; a model that still stumbles on 'not' is one worth double-checking before it filters a resume, a contract clause, or a medical symptom list.