The correction
The supplied research packet flags a specific reporting error: treating METR's February 2026 design update as though it withdrew the earlier developer slowdown result. The primary post does not do that. It restates the early-2025 finding and explains why a newer experiment became difficult to interpret. The important change is in the quality of the follow-up evidence, not in the historical status of the first result.
METR describes a selection problem: as AI became more useful and normal in participants' work, some developers were less willing to accept tasks that required working without it. If the people willing to participate differ from those who decline, the comparison can shift for reasons unrelated to the tool's effect. A more favorable-looking number from that setting does not automatically overturn the randomized earlier result.
How to read a productivity update
Keep three questions separate. What did the original experiment find? What changed in the tools and participants? How confident is the new estimate? An update can increase confidence that tools are improving while reducing confidence that the current experiment measures their effect cleanly. Those statements are compatible.
A team making adoption decisions should resist both easy headlines. 'AI always slows developers' outruns the 2025 task setting. 'The slowdown was retracted' misstates the 2026 post. The practical choice is to define a measurement window, include review effort, and note when participation or task selection makes a comparison unreliable.
- Read the first study and the update as separate documents.
- Record design changes and who opted into the newer experiment.
- Keep measured results separate from researchers' expectations about newer tools.
The friction to budget for
It is tempting to collapse changing evidence into a single directional verdict. That makes a short headline and a poor operating rule. The packet's correction is useful because it shows how even an evidence-focused summary can invert a finding by merging two different studies or time periods.
For workflow decisions, keep a dated evidence note: the population, task, tool period, result, and caveat. Revise the decision when stronger evidence arrives. Until then, uncertainty belongs in the conclusion rather than being hidden in a footnote.
The update is also a reminder that a tool can become attractive enough to change who agrees to a study. That is meaningful information about use, yet it complicates causal measurement. The two observations should be reported side by side rather than forced into a single percentage.
Sources & dates
- We are Changing our Developer Productivity Experiment Design ↗METR · Source date: 2026-02-24 · Retrieved 2026-09-16
METR restates its earlier slowdown finding while changing the design of a follow-up experiment. Developer reluctance to work without AI makes newer data an unreliable signal of the effect. The follow-up offers only weak evidence about a newer effect.
Source publication dates and historical event dates describe the source material. This page has not been publicly published.
