AI/ ai-benchmarks · llm-evaluation · legal-tech

Study Finds Legal AI Test Answers Guessable From Options Alone

A study of a 12,000-item Ukrainian judges exam shows top AI models can guess right answers from the options alone, exposing a benchmark design flaw.

A widely cited legal-exam benchmark for AI models turns out to be partly solvable without ever reading the question.

Researchers tested UA-JudgeExam, an 11,990-question, four-option benchmark built from Ukraine's judicial qualification exam, by hiding the question and showing AI models only the answer choices. Claude Haiku 4.5 still scored well above the 25% chance rate, and for 11.8% of items it picked the right answer no matter which of eight tested orderings the choices appeared in, a pattern pure chance would produce in about 0.2 items. A search across 280,059 editions of Ukrainian legislation ruled out the model simply recognizing quoted statute text, so the tell is something about how the correct answer is phrased relative to the wrong ones. Removing those obviously leaky items left 8,128 questions on which the gating model dropped to near-chance, but GPT-5.6, which had no role in building the filter, still answered 51.5% of them correctly with no question in sight.

That matters because raw scores on this kind of blind test are misleading on their own: Llama 3.1 8B posted a 29.2% blind score, beating every model except two, simply because it answers "A" on 92% of items and A happens to be a common key. Once researchers subtracted each model's own letter-picking habit, only GPT-5.6 and Claude Sonnet 4.6 kept a real edge, meaning most of the twelve models tested show no genuine answer-only leakage at all.

The fix isn't simple either: rewriting the wrong answers overcorrected into a new exploitable pattern, a smaller 400-item sample hid the problem completely, and a comparable exam benchmark, LEXam, showed no leak at all, suggesting this is a flaw in how these particular questions were written rather than something baked into legal exams generally.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →