AI/ legal-tech · rag · llm-evaluation · hallucination

New Benchmark Catches Legal AI Search Tools Making Things Up

A new bilingual dataset grades legal RAG systems claim by claim, and even top performers stumble on both retrieval and generation.

A new benchmark says the retrieval-augmented search tools law firms are starting to trust still invent things, and it can now show exactly where.

Researchers introduced ClaimRAG-LAW, a dataset built to stress-test retrieval-augmented generation (RAG) systems in legal search. Unlike prior legal RAG benchmarks, it covers both French and English, includes questions written for non-experts as well as lawyers, and spans a range of realistic query types. The team then ran current legal RAG systems through a fine-grained evaluation framework that scores retrieval accuracy, generation quality, and individual claims separately, rather than judging a system's output as one pass-or-fail block. The results showed measurable weaknesses in all three areas.

This matters because RAG is the industry's default answer to hallucination risk in high-stakes fields, and law is about as high-stakes as text generation gets. Most existing legal benchmarks are English-only and written for people who already know the vocabulary, which quietly ignores how non-lawyers actually use these tools. A benchmark that separates "the system fetched the wrong case" from "the system fetched the right case but described it wrong" gives builders a real diagnosis instead of a vague failure grade.

Call it an audit, not an indictment: the paper does not name which vendors failed where. But it confirms what plenty of legal-tech skeptics already suspected -- "RAG-powered" is not the same as "fact-checked."

TR

The Revision

Written by an AI system from the public sources credited above. How we write →