AI/ ai · benchmarks · forecasting · research

New Audit Finds AI Progress Forecasts Rest on Thin Data

A new audit of AI benchmark and compute data finds forecasts of frontier AI progress often rest on incomplete, unevenly sourced evidence.

A new academic audit finds that many confident predictions about frontier AI progress are built on a much thinner data foundation than the confidence suggests.

Researchers assembled a frozen record running through August 12, 2026, covering 62 AI systems, 12 versioned benchmarks, seven capability criteria, 144 graded events, 27 source records, and 408 typed relations connecting them. The gaps are stark: training compute figures are missing for 19 of 27 closed systems, including every closed model released so far in 2026, and none of the 35 open-weight systems in the sample report a METR task-horizon score. Only seven systems report both numbers together, the minimum needed to actually plot compute against capability. Where comparisons are possible, they're shaky - a seven-system link between two versions of the METR benchmark produced a slope of 1.206, but the study's statistical power is only strong enough to reliably catch swings of about 25 percent or larger.

The bigger problem is provenance: 73.2 percent of the 71 substantive quantitative events in the record trace back to a single measurement programme, and 76.1 percent come from lab-published releases rather than independent testing. That's a narrow, self-interested base for the calendar-dated milestone predictions that circulate around AI progress. The authors checked 56 other methodological sources and found 16 possible complementary measures, but no single number ready to replace the current patchwork.

So next time a clean trend line claims to predict when AI hits some milestone, ask whose numbers built it - there's a decent chance it's mostly one lab's.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →