Retrieval-augmented generation doesn't automatically make a language model smarter, and a new benchmark shows exactly when it backfires.
Researchers built a benchmark of 1,891 samples across five datasets and three task categories, then tested 11 LLMs using both sparse and dense retrieval methods. They measured three things: whether RAG beats skipping retrieval entirely, whether adding more documents helps, and whether the order of those documents changes the answer. Overall, models handled retrieval reasonably well, but robustness swung widely depending on the task. Qwen and GPT models lost significant ground when their reasoning mode was switched off, even on simple single-hop questions.
The results undercut the assumption that bolting retrieval onto any LLM is a free upgrade. Feeding retrieved documents to models as tool responses improved Claude's performance, but the same approach hurt Qwen and GPT, meaning a RAG pipeline tuned for one model family can't just be swapped onto another.
That's a more useful warning than another leaderboard score: RAG isn't a switch to flip on, it's a setup that needs testing per model and per task.