AI/ ai · benchmarks · nlp · model-evaluation

Study Finds AI Classifiers Struggle to Admit They Don't Know

A new benchmark finds AI text classifiers often miss when the right answer isn't offered, or wrongly reject valid answers instead.

A new benchmark shows two AI text classifiers are bad at admitting when the right answer isn't even on the list in front of them.

Researchers tested two classification models, labeled Laya and Jev, on four standard text-classification tasks: AG News, DBpedia, Emotion, and TREC, producing 23,932 predictions each from 300 calibration and 589 test texts. The core test removed the correct answer from a multiple-choice list and checked whether each model flagged it as missing rather than guessing. On TREC's five-option questions, Laya caught 97.2% of these missing-answer cases but wrongly rejected 69.7% of normal, fully-answerable questions, while Jev caught only 24.8% of missing answers but almost never rejected valid ones, at 0.0%. Adjusting the rejection threshold shifted the numbers to 33.9% detection and 3.7% false rejection for Laya, and 45.0% and 1.8% for Jev.

That's a real tradeoff, not a rounding error. A classifier that flags two-thirds of ordinary, answerable questions as unanswerable is useless in production, even if it never misses a genuinely missing answer. The paper's broader point, that accuracy, confidence ranking, and rejection behavior need to be measured separately, matters for anyone building AI systems that pick from fixed menus of answers.

The authors call this descriptive, not a fix, since it only covers deliberately removed reference labels, not genuinely open-ended questions, and offers no new rejection method, worth remembering before any vendor claims a model knows what it doesn't know.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →