METR's August 2025 developer productivity experiment yields an unreliable signal of AI's current effect on productivity.
Assessment
Evidence favors the claim, but the chain is incomplete or the sources are secondary.
In August 2025, METR began a follow-up to its early-2025 randomized experiment on AI tools and developer productivity, aiming to track how the effect was changing over time. In a February 2026 update, the experimenters themselves concluded that the new data gives an unreliable signal of AI's current productivity effect. The primary problem is selection: a significant and growing share of developers declined to participate because they did not want to work without AI, and this kind of opt-out by AI-reliant developers biases the measured speedup downward. METR also reported selection pressure from a reduced pay rate and unreliable time measurements on a fraction of tasks, and the resulting estimates carry confidence intervals spanning zero. Nor can participant surveys substitute for the compromised measurements, since developers' self-reported speedup estimates have themselves proven unreliable.
The unreliability is a matter of degree, not a total loss of information: METR notes the raw results show some evidence of speedup, and because the main known bias runs downward, the estimates can be read as conservative. But as a measure of the size of AI's current effect, the data cannot bear weight, a reading the experimenters state plainly and that independent commentary has accepted without credible dissent.
Full reasoning: the evidence and decisions behind this verdict
The claim originates with the experimenters. METR's February 2026 update (metr.org/blog/2026-02-24-uplift-update/) states directly: "we believe that the data from our new experiment gives us an unreliable signal of the current productivity effect of AI tools," and gives as the primary reason "a significant increase in developers choosing not to participate in the study because they do not wish to work without AI, which likely biases downwards our estimate of AI-assisted speedup." The post adds two further problems: selection effects from cutting pay from $150/hr to $50/hr, and unreliable time-on-task measurements for a fraction of tasks. The statistics are consistent with this self-assessment: returning developers show an estimated 18% speedup with a confidence interval of -38% to +9%, newly recruited developers 4% with -15% to +9%, both spanning zero.
The two supporting subclaims weigh as follows. Selection effects from AI-reliant developers opting out bias speedup estimates downward carries most of the load: it is METR's stated primary reason. It is not strictly required, since the pay-rate selection and measurement problems would leave the data noisy even without it. The unreliability of developers' self-reported speedup estimates closes the fallback route: if self-reports were trustworthy, surveys could recover the signal the measurements lost; METR's own early-2025 finding that slowed-down developers believed they had been sped up rules that out.
An adversarial check looked for credible sources arguing the data is in fact a reliable signal, including from skeptics of AI uplift who might suspect METR of walking back a null result. None was found: commentary (Rob Bowley's 2026 review, ScienceBlog's coverage) uniformly accepts the unreliability characterization, and the recorded instances all affirm. The strongest counter-reading, that the data remains informative as a conservative lower bound given the known bias direction, qualifies the claim without contradicting it, and is reflected in the assessment.
The verdict stops at supported rather than verified because the case rests substantially on METR's own report of participant feedback and surveys rather than on independently examinable data; the underlying participation records are not public. What would change the conclusion: a reanalysis showing the opt-out pattern was small or uncorrelated with AI reliance, or evidence that the confidence intervals were narrower than reported.
Decomposition
The claims this one rests on directly. ↗︎ opens a subclaim; the map shows how they fit together.
The claims this one rests on directly, not gathered into a named line of reasoning.
- supportsthis provides evidence for the parentsteward instructions →Selection effects from AI-reliant developers opting out bias measured AI productivity speedup estimates downward. ↗︎
- supportsthis provides evidence for the parentsteward instructions →Developers' self-reported estimates of AI-driven productivity speedup are unreliable. ↗︎
Provenance
Where this claim has been said, linked to its canonical form.
we believe that the data from our new experiment gives us an unreliable signal of the current productivity effect of AI tools
Discussion of the second, August 2025 developer productivity experiment.
They started a new experiment to track how things were changing – but couldn't complete it. They say the data was too compromised to produce reliable results. The interesting part is why the study broke down. Developers are now so reliant on AI that they won't work without it
A commentary reviewing METR's 2026 research update; the author endorses in his own voice that the follow-up experiment broke down and could not produce reliable results, attributing the compromise to developers' unwillingness to work without AI.
Cite this claim: a formal citation with its evidence attached
Contribute
Every judgment on this page is open to challenge. A contribution is evaluated on its merits by the reviewer; if it succeeds the page changes, and if it does not, the reasons are stated. Either way the exchange becomes part of the claim’s public record.
The attention this claim received was paid for by a funded mandate. Funding buys only scheduling: it can make an assessment happen sooner, or reach deeper into a subtree. It has no influence on what the assessment concludes, and none on which claims enter the graph; assessments run under the same public standards whoever pays, funders never see or shape a verdict before anyone else, and mandates that attempt to steer conclusions are refused.
Created by extractor · Aug 10, 2026. Every judgment on this page is accompanied by a reasoning trace.