Historical source and event dates are not site publication dates. Product plans, policies and availability may have changed since retrieval.

The setup
METR is an AI-evaluation research organization, and its March 2025 paper set out to measure something narrower than a general capability score: the length of a task, timed by how long a skilled human would take, that a leading AI agent can complete on its own with a given reliability. The original publication built this metric from a suite of software and research tasks, including the HCAST and RE-Bench sets and a comparison against SWE-Bench Verified, each scored by whether an agent's attempt succeeded, not by a human judge's impression.
What the documents show
METR's paper reports that the length of task a frontier model agent could complete with 50 percent reliability had been doubling roughly every seven months over six years, and that extrapolating this trend implies agents able to independently finish tasks currently taking people days or weeks in under a decade. METR frames this as a measured historical trend, not a guarantee; the paper is explicit the projection is an extrapolation. METR's separate, continuously updated time-horizons page, last revised May 2026 at the time this note was prepared, carries the same metric forward for newer models on a larger task suite, describing itself as a living successor to the original paper's fixed dataset.
The friction
METR's own paper names the limits of what its metric captures: a plot of the trend does not account for future changes or for external validity concerns, which it says are responsible for the majority of the uncertainty in its forecast, more than the statistical noise in the tasks themselves. The tasks are well-defined problems with a checkable success condition, which is what makes them scoreable, and that same property is what makes them different from the ambiguous, priority-shifting work a real job usually involves. METR discusses this directly, noting most benchmarks fail to cover a wide enough difficulty range for this kind of measurement, part of why it built a new one.
What changed in the work
What the paper supports is a checkable claim about a specific set of scored tasks: newer models complete longer scored tasks than older ones, at a fairly consistent rate. It does not support a claim that a team's actual workload, with its ambiguity and shifting requirements, is shrinking on the same schedule; that would extend the benchmark's result past what METR's paper claims. Reading the trend as a ceiling on what is technically possible in narrow, well-specified work, rather than a forecast of general workplace productivity, is this note's editorial framing of where METR's caveats leave the result.
- Does the team's task resemble METR's checkably-scored software tasks, or the messy work METR says its metric does not capture?
- Has the team checked METR's current time-horizon figures rather than the March 2025 numbers?
- What would external validity failing look like for the task the team is trying to extrapolate to?
METR's paper earns its citation by being candid about its own metric's limits, naming external validity as its largest uncertainty, precisely the caveat most likely to get dropped when the seven-month doubling figure travels without its source attached.
Sources & verification
Preserved from the earlier archive. These sources have not all been freshly rechecked for this expansion.
- Measuring AI Ability to Complete Long Software TasksSource date: 2025-03-19 · Retrieved: 2026-09-16
States the length of software task an AI agent can complete at 50% reliability has doubled roughly every 7 months over 6 years, and names external validity as the largest source of uncertainty in extrapolating the trend.
- Task-Completion Time Horizons of Frontier AI ModelsSource date: not stated · Retrieved: 2026-09-16
Living tracking page giving the current 50%- and 80%-time-horizon measurements for frontier models on a larger task suite, last updated May 2026.
Continue the workflow
- Run a workflow pilot that can answer a real question
Decide whether a proposed workflow deserves wider use using a modest, honest pilot.
- Epoch AI's compute numbers are mostly estimates, and it says so
Epoch AI's own methodology shows most training-compute figures in its tracker are estimated, not measured, and it labels the difference.
- A model maker's own usage index counts conversations, not the economy
Anthropic's first Economic Index sorts a million Claude.ai conversations by occupation but says it cannot show AI use off its platform.
- Stanford's AI Index reports numbers it credits to other sources
The annual AI Index compiles labor-market and economy figures from named outside providers rather than measuring them itself.