AI/ ai-benchmarks · forecasting · llm-evaluation · data-contamination

New Benchmark Tackles AI Forecasting's Leakage Problem

LEAF, a new living benchmark, slashes future-information leakage in AI forecasting evaluations from 8.6% to 1.6% and tests whether 16 models can forecast.

A new benchmark says most AI forecasting tests have been cheating without telling anyone.

Researchers built LEAF, a benchmark for testing how well large language models predict trends, events, and time series when fed outside news and data feeds. The core problem: automated web searches during evaluation often surface articles published after the event being predicted, letting a model "forecast" something it already read about. LEAF's fix is a two-agent system - one agent retrieves context, a second cross-checks it for future leaks - paired with a recursive retrieval setup meant to keep auxiliary material temporally honest. The team says a manual audit of 500 tasks by 47 domain specialists found this pipeline cut leakage from 8.6% of cases down to 1.6%.

Contaminated benchmarks are a quiet problem in AI evaluation - a model that scores well because it peeked at the answer looks identical, on a leaderboard, to one that actually reasoned its way there. By testing 16 proprietary and open-weight models and finding they genuinely extract useful signal from verified, leak-checked events, LEAF offers a cleaner read on whether LLMs can do real forecasting rather than retrieval with extra steps.

"Living" is the operative word here: if LEAF's event pool doesn't keep refreshing, it risks becoming exactly the kind of stale, memorizable test it was built to replace.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →