Security/ ai security · rag · backdoor attacks · arxiv research

A Poisoned Retriever Can Quietly Hijack AI Search Agents

A new arXiv paper shows a backdoored retriever can hide malicious behavior by faking a security fix, without touching the underlying data.

A single bad model checkpoint can turn a trustworthy AI search agent into a puppet, and hide the strings before anyone looks.

That is the finding of a September 30, 2026 paper posted to arXiv (arXiv:2609.37468), "Backdoor in the Loop: Compromising Agentic Search via Malicious Retrievers." The researchers focus on agentic retrieval-augmented generation, where an AI agent repeatedly queries a retriever and lets the results shape its next search. They show an attacker who supplies only a poisoned retriever checkpoint - never touching the underlying document corpus - can still suppress useful evidence, keep steering the agent back to one planted document, or drag out searches to run up retrieval, context, and latency costs. To dodge detection, the attacker plants a second, weaker backdoor and then "unlearns" it, producing a false signal of purification while the original backdoor stays intact.

That last trick is the real headline. Most RAG security advice focuses on locking down the corpus - the documents an agent can retrieve. This paper argues the retriever model itself is a separate, under-guarded attack surface, and that today's backdoor detectors can be actively gamed rather than just fooled by chance. For any team treating a clean corpus as equivalent to safe search, that is a gap worth closing.

It is a familiar shape wearing new clothes: a software supply-chain attack, translated into a model checkpoint. Enterprises learned to vet dependencies after years of poisoned npm and PyPI packages; agentic AI pipelines are about to relearn that lesson with retrievers instead of libraries.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →