METR's headline AI benchmark just got a math audit, and the scale underneath it turns out to be lumpier than advertised.
Researchers reanalyzed METR's 50% time horizon metric, which scores an AI by the length of human-coded tasks it can complete with 50% success, across 228 tasks and 26 models. The original approach assumed AI difficulty scales linearly with the log of human completion time. The new analysis instead fits splines and applies item-response theory to let the data define the curve. The result is nearly flat for tasks that take a person 2 to 30 minutes, then climbs close to linearly beyond that, meaning a jump from a 3-minute task to a 30-minute one is far less demanding for an AI than the same 10x jump from 30 minutes to 5 hours.
Time horizons get cited as shorthand for how fast AI is closing in on human-level autonomy on real work, often implying steady, even progress. If the underlying scale is uneven, comparing jumps across models or benchmark versions can overstate or understate real gains, a risk that grows as benchmarks add longer tasks. The authors argue the metric needs company: diagnostic plots showing where the curve bends, not just a single headline number.
A number that looks like a clean doubling can just be the ruler getting weird, not the AI getting twice as capable.