Desk Trials

CODING PRODUCTIVITY EVIDENCE / 2 MIN READ

What the developer slowdown study actually measured

A surprising result matters most when its task and participants stay attached.

The evidence

The headline result is memorable: experienced open-source developers took roughly 19% longer with AI tools on tasks in repositories they knew well. METR's July 2025 study is narrower, and therefore more useful, than a blanket claim that coding assistants slow developers down. It involved 16 developers and 246 tasks in mature projects. These were people working in familiar code, with the AI available as an option.

That setup gives the result bite. The developers were not simply unfamiliar beginners who had never prompted a model. Yet it also limits where the result applies. A new hire in an unknown codebase, a disposable prototype, or a repetitive migration may face different constraints. The study's own discussion warns against treating its finding as a verdict on all software work.

A proposed team check

If a team wants to know whether coding assistance helps, choose a few representative tasks and define what completion includes: implementation, review, tests, and cleanup. Record elapsed effort, rework, and defects. Separate familiar-repository changes from new or exploratory work. Then compare the whole task, not just the first generated patch.

The setup cost is real. Someone must identify comparable tasks and avoid letting volunteers select only the work they expect AI to improve. Perception is also worth recording, but it should not replace elapsed time or outcome quality. METR reported a gap between developers' speed estimates and measured completion in this setting. That gap is a reason to keep both measures visible.

  1. Name the task type and developer experience before measuring.
  2. Count review and repair time in completion.
  3. Report results by context instead of one organization-wide average.

The practical limit

A single study from early-2025 tools does not freeze the category forever. Tool capability, use patterns, and participation can change. The more durable lesson is methodological: a persuasive demo and a pleasant editing session are not measurements of shipped work.

The result worth keeping is a local decision about where assistance earns its review cost. For some tasks that may be most of the workflow; for others the best use may be a narrow explanation or test suggestion. Keep the task boundary attached to every claim about speed.

If a pilot reports that developers like the tool, report that as satisfaction. If it reports shorter completion, explain what was timed. If it reports fewer defects, show the review period. Those are different outcomes, and collapsing them into one productivity score makes the finding less useful.

Sources & dates

  1. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity ↗METR · Source date: 2025-07-10 · Retrieved 2026-09-16

    The randomized study involved 16 experienced open-source developers and 246 tasks in mature projects. Developers using early-2025 AI tools took about 19% longer in this setting. METR cautions against generalizing the result to all software engineering.

Source publication dates and historical event dates describe the source material. This page has not been publicly published.

Keep reading