A handful of booby-trapped documents can hijack what a RAG-powered chatbot tells you.
Researchers describe an attack called BadRAG that targets retrieval-augmented generation systems, the setup where a chatbot pulls in text from an external knowledge base before answering. Many of those knowledge bases are large, unsanitized pools of user-generated content, the kind of thing Google Search draws from when it surfaces sites like Reddit. In the attack, someone slips a small number of malicious passages into that data, built in two stages: first to be retrieved only when a user's query contains a specific attacker-chosen trigger word, then to push the model toward a bad outcome once retrieved, whether that is refusing to answer, flipping sentiment, leaking hidden context, or misusing a tool the model has access to. None of this requires touching the user's question or retraining the model itself.
The scale is the unsettling part. Just 10 malicious passages, 0.04% of the test corpus, pushed retrieval success to 98.2% and drove negative responses from a baseline of 0.22% up to 72% whenever a trigger word appeared. That is a tiny, cheap payload for a large, targeted effect on any product that quietly ingests public text, which describes most RAG deployments running today.
Prompt injection got the headlines this year. BadRAG is a reminder that the attack surface for LLMs is not just what users type in, it is whatever the model is allowed to go read.