A team of researchers has built a scoring model that judges creativity in debate more accurately than off-the-shelf large language models.
The paper introduces DEFINED, a framework that operationalizes creativity through an eight-dimensional metric system covering both divergent thinking and convergent thinking. The model was trained on transcripts and expert scores pulled from real debate competitions - not synthetic benchmarks or crowdsourced labels. One practical problem the researchers had to solve: publicly available debate data skews heavily toward elite performers. They applied a constrained data augmentation strategy to broaden coverage toward mid-to-low proficiency debaters. Across their evaluation protocol, the specialized model outperformed prompt-based LLM evaluators and existing debate scoring methods.
Debate is a reasonable testbed for this kind of work - it is data-rich, publicly accessible, and captures creativity under real pressure in a way that standard divergent-thinking tasks never do. More broadly, the result adds to a growing body of evidence that general-purpose LLMs can be beaten on specific evaluation tasks by narrower models fine-tuned on domain data. That matters most as AI-judged output becomes the default in education, hiring, and research.
Whether an eight-dimensional rubric actually captures what humans mean by creativity is a question the paper approaches carefully but does not fully settle - which, to be fair, is true of every creativity framework that has come before it.