A fine-tuned, domain-specific language model just outclassed GPT-3.5 and GPT-4 at a real legal classification task - and it didn't need anywhere near their size to do it.
Researchers pitted a range of text-classification approaches, from traditional machine learning to large language models, against ten categories of Korean sexual offense court precedents. KLUE-BERT, a Korean-language model fine-tuned specifically on the legal text, hit 99.3% accuracy, beating both GPT-3.5 and GPT-4 as well as the traditional machine learning baselines. The team then used explainable AI (XAI) techniques to pick apart why the models got things right or wrong, identifying which linguistic features drove predictions. When they tested the fine-tuned model on KICS data, which more closely resembles messy real-world case records, it stumbled on implicit contextual cues that a human reader would catch.
That gap matters more than the headline accuracy number. It suggests domain-specific fine-tuning beats throwing a bigger general-purpose model at a specialized legal task, which should temper any assumption that GPT-4-class tools are automatically the right choice for legal document review. It also matters for anyone trying to deploy this kind of system in a courtroom or law firm: knowing exactly where a model's judgment breaks down, via XAI, is the difference between a useful assistant and a liability.
A model that aces curated precedent categories but trips on messier real-world phrasing isn't ready to replace a paralegal yet. It's a reminder that fine-tuning is not the same as understanding, and that legal AI's hardest problem is still reading between the lines.