A new benchmark says most AI forecasting tests have been cheating without telling anyone.
Researchers built LEAF, a benchmark for testing how well large language models predict trends, events, and time series when fed outside news and data feeds. The core problem: automated web searches during evaluation often surface articles published after the event being predicted, letting a model "forecast" something it already read about. LEAF's fix is a two-agent system - one agent retrieves context, a second cross-checks it for future leaks - paired with a recursive retrieval setup meant to keep auxiliary material temporally honest. The team says a manual audit of 500 tasks by 47 domain specialists found this pipeline cut leakage from 8.6% of cases down to 1.6%.
Contaminated benchmarks are a quiet problem in AI evaluation - a model that scores well because it peeked at the answer looks identical, on a leaderboard, to one that actually reasoned its way there. By testing 16 proprietary and open-weight models and finding they genuinely extract useful signal from verified, leak-checked events, LEAF offers a cleaner read on whether LLMs can do real forecasting rather than retrieval with extra steps.
"Living" is the operative word here: if LEAF's event pool doesn't keep refreshing, it risks becoming exactly the kind of stale, memorizable test it was built to replace.