Researchers have built a test that catches language models contradicting themselves, plus a training method that stops it.
A team working with Qwen2.5 and Phi-3.5 models borrowed an idea from betting markets: if a model's probability estimates on logically related claims let a trader make guaranteed money no matter what happens, the model doesn't understand what it's saying. The researchers call this a Dutch book, and treat the amount of exploitable inconsistency as a measurable score rather than a yes-or-no judgment. They also show, mathematically, that the standard way models are trained, predicting the next word, produces this incoherence even at its theoretical best, so the flaw is baked into the training objective, not just sloppy execution. Their fix, a framework called Arbitr, adds an adversarial trader during training that penalizes the model whenever its answers can be exploited, plus a calibration anchor so the model can't dodge bets by just going vague.
This matters because most AI evaluations test whether an answer is factually right, not whether a model's beliefs hang together across differently worded versions of the same question, a blind spot distinct from the hallucination benchmarks that dominate current AI research. Arbitr cut that exploitability by orders of magnitude in the paper's five pre-registered experiments without sacrificing task accuracy, and the effect carried over to logical patterns and model families it never saw during training.
The catch undercuts the sales pitch: at 7 billion parameters, the models that looked most logically consistent were also the most overconfident, proof that making a model harder to out-bet is not the same as making it right.