AI/ ai agents · legal tech · benchmarks · llm evaluation

AI Legal Research Agents Still Get Most Cases Wrong

A new benchmark shows even the best legal research AI agent, Claude Opus 4.8, nails only 42.9 percent of complex legal research questions.

The best AI legal research agent is right less than half the time.

Researchers built Legal Research Bench, a set of 413 expert-written U.S. legal research questions, each paired with a gold answer, supporting case law, and a strict pass-fail rubric. They tested thirteen frontier models in an agent harness equipped with web search, case-law search, and page-parsing tools, then checked every cited authority for accuracy rather than just grading the final answer. Claude Opus 4.8 topped the field but still answered only 42.9% of questions fully correctly. The team also validated its grading against practicing attorneys to confirm the scores track real legal judgment, not just a quirk of the rubric.

Legal research is exactly the kind of long, citation-heavy task AI agents are supposed to be good at - lots of documents, lots of retrieval, a clear right answer. But the study found that giving models more turns, more tool calls, or more inference cost did not make them more accurate, and scores dropped further on questions requiring reconciliation of conflicting authorities.

A model that misses a stale citation isn't just wrong - it can get a lawyer sanctioned. Fewer than half right is a long way from an AI paralegal you'd let touch a real case unsupervised.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →