AI/ ai · benchmarks · llm-evaluation · research

Benchmark Buries the Question in the Puzzle, Not the Prompt

EnigmaForge hides its question in fictional documents, and asking AI to find it first scrambles the usual model rankings.

A new benchmark stops telling AI models what the question is.

EnigmaForge, detailed in a new arXiv paper, drops models into a pile of fictional letters, receipts, and logbook margins that hide a small logic puzzle with exactly one solution, verified by a SAT solver, with a certificate confirming every clue actually matters. Because the puzzles are generated rather than hand curated, the test set can keep growing forever. Researchers ran 25 frontier models across 600 instances, producing 17,400 scored records, under three matched conditions that measured both raw "intuition" (solving with no question given at all) and simple fact recovery from the same documents.

The two scores tell very different stories. Models spread by 22x on intuition but only 1.6x on fact recovery, and the model ranked second best at pulling out facts falls to fourteenth once it has to notice the puzzle on its own. That gap is the real finding: benchmarks built on models answering questions they are handed may say little about whether a model can spot that there is a question worth answering in the first place.

Some models never got that far. Their own content filters blocked them before they reached the puzzle, a reminder that any benchmark counting refusals as failures is quietly grading filter behavior instead of reasoning.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →