AI chatbots can narrate their chess moves fluently. A new study finds that narration often has little to do with why the move was actually picked.
Researchers tested explanations across 200 Lichess endgame puzzles using three checks: move recoverability, decoder-side controls, and token-level scoring of legal candidate moves. When explanations still contained explicit move hints, it was easy to recover the intended move from the text alone. Strip those hints out, though, and the advantage nearly disappeared, leaving explanations that added only small, decoder-dependent gains over just handing a model the bare board state. Token-level scoring turned up something odder still: feeding a model a plausible-sounding explanation lifted from an unrelated puzzle actually lowered the probability it assigned to the correct move, and recognizable endgame motifs made moves easier to recover without making the models any more likely to play correctly.
That distinction matters well beyond chess. "The model explained its reasoning" is increasingly treated as evidence that the reasoning can be trusted, in domains far messier than an endgame puzzle. Chess works as a clean test bed precisely because the board is fully observable and move quality is objectively checkable - most real-world explanation demands offer no such luxury.
If an AI can be nudged by a stranger's irrelevant explanation borrowed from a different puzzle, "explain your reasoning" may be measuring narrative fluency, not faithfulness.