A new benchmark grades AI models on curiosity, not correctness.
CruxBench, introduced in an arXiv paper published today, tests whether large language models can identify the right subquestions - the researchers call them "cruxes" - on the way to solving a hard forecasting problem. Rather than checking a model's final answer against a fixed label, it scores each proposed question by its Value of Information, or how much that question would actually shift beliefs about a real future event. The researchers ran eight models against 293 forecasting questions. VOI scores correlated strongly with independent measures of model capability (r=0.90), but even frontier models only narrowly beat a baseline that ignores question content entirely and guesses based on timing.
Most AI benchmarks measure whether a model can answer a question someone else already framed for it. CruxBench flips that setup: it tests whether a model can figure out which question is worth asking, a skill closer to how research, due diligence, and strategic planning actually work in practice. That top models barely edge out a timing-only baseline suggests genuine information-seeking is not something scale has quietly solved along the way - a real capability gap, not a footnote.
The design also has a built-in defense against the usual complaint about AI benchmarks: contamination. Ground truth here does not exist until the future actually happens, so no model can have memorized the answer key.