A new benchmark says the retrieval-augmented search tools law firms are starting to trust still invent things, and it can now show exactly where.
Researchers introduced ClaimRAG-LAW, a dataset built to stress-test retrieval-augmented generation (RAG) systems in legal search. Unlike prior legal RAG benchmarks, it covers both French and English, includes questions written for non-experts as well as lawyers, and spans a range of realistic query types. The team then ran current legal RAG systems through a fine-grained evaluation framework that scores retrieval accuracy, generation quality, and individual claims separately, rather than judging a system's output as one pass-or-fail block. The results showed measurable weaknesses in all three areas.
This matters because RAG is the industry's default answer to hallucination risk in high-stakes fields, and law is about as high-stakes as text generation gets. Most existing legal benchmarks are English-only and written for people who already know the vocabulary, which quietly ignores how non-lawyers actually use these tools. A benchmark that separates "the system fetched the wrong case" from "the system fetched the right case but described it wrong" gives builders a real diagnosis instead of a vague failure grade.
Call it an audit, not an indictment: the paper does not name which vendors failed where. But it confirms what plenty of legal-tech skeptics already suspected -- "RAG-powered" is not the same as "fact-checked."