Security/ rag · llm-security · data-poisoning · ai-research

Researchers Show 10 Poisoned Passages Can Hijack a RAG Chatbot

A new attack called BadRAG shows that a handful of poisoned passages in a knowledge base can make a chatbot refuse to answer or turn hostile.

A handful of booby-trapped documents can hijack what a RAG-powered chatbot tells you.

Researchers describe an attack called BadRAG that targets retrieval-augmented generation systems, the setup where a chatbot pulls in text from an external knowledge base before answering. Many of those knowledge bases are large, unsanitized pools of user-generated content, the kind of thing Google Search draws from when it surfaces sites like Reddit. In the attack, someone slips a small number of malicious passages into that data, built in two stages: first to be retrieved only when a user's query contains a specific attacker-chosen trigger word, then to push the model toward a bad outcome once retrieved, whether that is refusing to answer, flipping sentiment, leaking hidden context, or misusing a tool the model has access to. None of this requires touching the user's question or retraining the model itself.

The scale is the unsettling part. Just 10 malicious passages, 0.04% of the test corpus, pushed retrieval success to 98.2% and drove negative responses from a baseline of 0.22% up to 72% whenever a trigger word appeared. That is a tiny, cheap payload for a large, targeted effect on any product that quietly ingests public text, which describes most RAG deployments running today.

Prompt injection got the headlines this year. BadRAG is a reminder that the attack surface for LLMs is not just what users type in, it is whatever the model is allowed to go read.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →