AI/ ai · llm-evaluation · classification · prompt-engineering

Study Finds AI Classifiers Follow Labels, Not Definitions

Researchers found open-weight AI classifiers often judge options by their label, not their definition, and one line of prompt code decides which.

New research shows some AI classifiers decide what an option means by its label, not by the rule a developer actually wrote for it.

The study examines typed decision models, an interface popularized by a system called Jev for moderation and routing tasks, in which a model scores several options that each carry a short label and a written definition. Researchers tested four open-weight typed decision models, three ways of reading answers from a Qwen2.5 backbone, eleven classification tasks, and a new synthetic suite they built called PolicyBench. In one system, laya-td, deleting every definition left accuracy almost unchanged (0.8559 versus 0.8487), even though the definitions alone scored 0.7971 on their own, and renaming the options to plain "A" and "B" lifted accuracy by about 15 points.

The researchers traced the bug to one line of code: laya's prompt template writes each option as "label: definition," while a comparable system called von renders only the definition. Flipping that single line, with no weights touched, made laya's bias disappear entirely and created the identical failure in von, whose accuracy collapsed from 0.8511 to 0.2281 once a label was made to contradict its definition. That contradicts earlier work blaming the model's constrained decision head - this is a prompt-rendering bug, not a model limitation.

Any team that assumed a carefully written policy was steering its moderation or routing system might actually be watching label names do the deciding, and the fix on offer here is one line of template code, not a retrain.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →