AI/ legal-ai · hallucination · benchmark · llm-evaluation

New Benchmark Shows AI Still Fakes Legal Citations

A new benchmark built from New York appeals court rulings finds top AI models still rate fabricated legal citations as valid, even with full case text.

AI legal assistants are good at sounding right. A new benchmark shows how often they are not.

Researchers built PARCEL, a benchmark that checks whether a legal claim is actually backed by the case it cites. Using 3,396 parenthetical-style claims pulled from recent New York State Court of Appeals decisions, they labeled each one Supported, Refuted, or Not Found, then tested several leading LLMs in a zero-shot setup. The strongest models hit up to 0.97 accuracy on the task. But even when given the full text of the court opinion, models still mislabeled unsupported claims as supported.

That failure mode matters more than the headline number. Across every model tested, spotting a missing citation was harder than catching a flat-out contradiction, and fabricated-but-plausible citations caused the biggest accuracy drop of all. In other words, the errors that are hardest to catch are exactly the ones a tired lawyer skimming a brief would miss too.

A 97% accuracy score reads well in a demo. It is less reassuring once you remember that the remaining 3% is where a citation that does not say what it claims to say slips straight into a legal filing.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →