AI/ rag · ai-benchmarks · compliance-ai · enterprise-ai

Generic AI Document Search Falls Short on Compliance Questions

A benchmark of 200 questions found generic upload-and-ask document retrieval trails a version- and scope-aware system by 9.6 points on regulatory documents.

Uploading a folder of regulations to a hosted AI search tool and asking questions is not as reliable as it looks.

A new study tested that "upload the documents and ask" approach against a governed system that explicitly checks a document's version, jurisdiction, subject, and date before it generates an answer. Both systems were evaluated on 200 questions drawn from a stratified sample of a published benchmark, each paired with a known correct source document, pulled from a pool of about 73,000 candidate normative documents used in a production deployment. The hosted retrieval service scored 88.1. The governed system scored 97.7, a gap of 9.6 points on unrounded means.

The gap matters because "which passage sounds relevant" and "which passage is the current, legally applicable one" are different questions, and generic hosted retrieval conflates them. For any product answering questions about rules that vary by jurisdiction and change over time, ignoring that distinction turns a wrong answer into a confident-sounding one.

The governed system is not a research prototype. It has run as a commercial product since January 2026, has 1,126 registered users, counts Zhipu AI and Lecheng Health as customers, and was handling roughly 100,000 calls a workday by mid-April 2026. That commercial stake is worth keeping in mind: this is as much a vendor showing its product beats the generic default as it is a neutral study, though the question set, answers, and scoring scripts are all public for anyone who wants to check the math.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →