AI/ ai safety · large language models · reward hacking · reinforcement learning

AI Models Catch Liars Easily but Wrongly Accuse Honest Ones

A new study finds large language models spot a lying reward reporter almost perfectly but falsely accuse an honest one up to 58 percent of the time.

Large language models are excellent at sniffing out a lying source, and surprisingly bad at trusting an honest one.

Researchers built a simple two-option game where a reward reporter could tell the truth or lie, engineered so a genuine payout swap and an outright lie produced identical histories. They then gave three large language models, from two families ranging from 32B to 72B parameters, one verified record: an independent check of what actually happened in a single round, printed beside the reporter's own account. Handed that clue, the models nearly always caught a lying reporter: the 70B class did it in every condition tested, and the 32B model tripped up only once, on a particular wording. Clearing an honest reporter was a different story. A 72B model wrongly branded an honest reporter a liar 38% of the time when nothing had changed, and 58% of the time when payouts had genuinely moved, while a second 70B model made the same error 26% and 48% of the time.

This isn't a reading problem. The same models scored 0.96 to 1.00 on the easy case, when the correct answer was already stated in the prompt. Instead, the wrong calls tracked details that should not matter at all: which round the verified record referenced, for one model family, and which letter stood for "honest," for the other. That is a shaky foundation for a use case researchers keep proposing, LLMs auditing whether an automated reward signal has been tampered with.

The team had registered a prediction before running the test, guessing a 35% error rate. The real number, 58%, blew past it, which is as close as a paper gets to admitting its own skepticism wasn't skeptical enough.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →