AI/ ai · ai benchmarks · reasoning · metacognition

Japanese riddles expose a blind spot in AI reasoning

A new benchmark built from kids' riddles shows top language models know the right answer more often than they're willing to say it.

A benchmark built from Japanese children's riddles just caught frontier AI models doing something odd: getting the right answer, then rejecting it.

Researchers assembled 201 nazonazo, a genre of Japanese wordplay riddles that require reinterpreting a question rather than recalling a fact, then tested 38 frontier language models released between 2023 and 2025 under retrieval-free, zero-shot conditions, meaning no web lookups and no practice on the exact questions. They compared the models against 126 human solvers who averaged 52.9% accuracy on a 120-item subset. Non-reasoning models scored 7.6%. Models built for step-by-step reasoning did better, but still only reached 17.6%. The more telling finding came from reading the models' own reasoning transcripts: models frequently produced the correct answer as a candidate partway through their reasoning, then abandoned it for a wrong final answer.

That gap between generating a good idea and trusting it is the real story here. The researchers call it verification failure, and say it accounts for between 5% and 39% of a given model's wrong answers, depending on which model you look at. It suggests the constraint isn't creativity, it's judgment: these systems can stumble onto insight but lack a reliable way to recognize it as insight.

That's a more useful diagnosis than another leaderboard, especially given how many reasoning benchmarks these models ace partly because the answers leaked into training data somewhere along the way. A puzzle set that humans still only solve about half the time, and that can be refreshed indefinitely, is a harder thing to game - which may be exactly why the scores here look so much worse.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →