AI/ ai-benchmarks · model-calibration · llm-evaluation

New Benchmark Exposes Flaws in AI Confidence Scoring

TypedBench finds AI confidence scores are wording-sensitive and underconfident, which can backfire once real costs enter the decision.

A new benchmark says the confidence scores behind automated AI decisions often cannot be trusted.

Researchers built TypedBench, a benchmark for so-called System One decision models - AI systems that output calibrated probabilities over typed answers like categories, ordinal levels, or yes/no outcomes, rather than generating text. The benchmark draws on seven policy-labelled generators and nine evaluation suites, testing accuracy across paraphrased wording, calibration error against a theoretical perfectly-calibrated predictor, and performance under asymmetric cost matrices. The team ran it on a hosted model, an open encoder, and a family of open decoders ranging from 0.8 billion to 9 billion parameters, all on identical test items.

The hosted model stuck to its stated policy, but its confidence scores shifted depending on how a question was phrased, and it was systematically underconfident. Under cost-sensitive conditions, that underconfidence was bad enough that just taking the model's top answer beat actually using its probabilities. The open decoders were exact but got slower as more options were added, and their accuracy dropped on policy questions as the choice set grew.

That is a real problem for anyone wiring these scores into automated triage, routing, or escalation systems: a probability that moves when you rephrase the question is not really a probability.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →