AI/ ai · research · nlp · evaluation

Purpose-Built Debate Scorer Outperforms General LLMs

A new framework trained on real debate competition data scores human creativity across eight dimensions, more accurately than prompt-based large language models.

A team of researchers has built a scoring model that judges creativity in debate more accurately than off-the-shelf large language models.

The paper introduces DEFINED, a framework that operationalizes creativity through an eight-dimensional metric system covering both divergent thinking and convergent thinking. The model was trained on transcripts and expert scores pulled from real debate competitions - not synthetic benchmarks or crowdsourced labels. One practical problem the researchers had to solve: publicly available debate data skews heavily toward elite performers. They applied a constrained data augmentation strategy to broaden coverage toward mid-to-low proficiency debaters. Across their evaluation protocol, the specialized model outperformed prompt-based LLM evaluators and existing debate scoring methods.

Debate is a reasonable testbed for this kind of work - it is data-rich, publicly accessible, and captures creativity under real pressure in a way that standard divergent-thinking tasks never do. More broadly, the result adds to a growing body of evidence that general-purpose LLMs can be beaten on specific evaluation tasks by narrower models fine-tuned on domain data. That matters most as AI-judged output becomes the default in education, hiring, and research.

Whether an eight-dimensional rubric actually captures what humans mean by creativity is a question the paper approaches carefully but does not fully settle - which, to be fair, is true of every creativity framework that has come before it.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →