Retrieval-augmented AI models will drop a correct answer the moment one search result disagrees with it.
Researchers built a diagnostic called RAG-Stress to test how easily retrieval-augmented generation systems abandon answers they already got right. They ran it on fifteen systems, including API models, open models, and search agents trained with reinforcement learning, across TriviaQA-RC, HotpotQA, SearchQA, and English and Chinese medical QA sets. Starting from questions a model answered correctly with no outside lookup, they edited one sentence in the retrieved evidence to support a specific wrong answer, then moved that doctored sentence to the start, middle, or end of the passage. Telling a model to prioritize the documents over its own knowledge raised the rate of wrong flips by 10.9 to 13.5 percentage points versus instructions allowing it to trust prior knowledge, and models were misled most often when the false claim sat at the end of the passage.
Standard accuracy scores can't separate a model correctly updating on new evidence from a model getting steamrolled by a single bad sentence, which makes retrieval-augmented systems look more reliable than this test suggests. A follow-up audit of 500 questions across two model checkpoints found harmful overrides increasing without a matching rise in cases where retrieval actually caught and fixed a wrong answer.
Retrieval augmentation was pitched as a fact-check layer on top of a language model's guesses. This study is a controlled demonstration that the layer can be talked out of the truth by one well-placed false sentence, no matter how confident the model was a moment earlier.