AI/ ai benchmarks · metr · ai research · time horizons

A Statistical Gut Check for AI's Favorite Capability Metric

A new statistical analysis of METR's AI time horizon benchmark finds the metric's scale is uneven, making some capability jumps look bigger than they are.

METR's headline AI benchmark just got a math audit, and the scale underneath it turns out to be lumpier than advertised.

Researchers reanalyzed METR's 50% time horizon metric, which scores an AI by the length of human-coded tasks it can complete with 50% success, across 228 tasks and 26 models. The original approach assumed AI difficulty scales linearly with the log of human completion time. The new analysis instead fits splines and applies item-response theory to let the data define the curve. The result is nearly flat for tasks that take a person 2 to 30 minutes, then climbs close to linearly beyond that, meaning a jump from a 3-minute task to a 30-minute one is far less demanding for an AI than the same 10x jump from 30 minutes to 5 hours.

Time horizons get cited as shorthand for how fast AI is closing in on human-level autonomy on real work, often implying steady, even progress. If the underlying scale is uneven, comparing jumps across models or benchmark versions can overstate or understate real gains, a risk that grows as benchmarks add longer tasks. The authors argue the metric needs company: diagnostic plots showing where the curve bends, not just a single headline number.

A number that looks like a clean doubling can just be the ruler getting weird, not the AI getting twice as capable.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →