Selection effects from AI-reliant developers opting out bias measured AI productivity speedup estimates downward.
Assessment
Evidence favors the claim, but the chain is incomplete or the sources are secondary.
The claim originates with METR, whose February 2026 update to its developer productivity experiment reported that the primary problem with its new data was a significant rise in developers declining to participate because they did not wish to work without AI, with commentators citing refusal rates of 30 to 50 percent. That behavioral fact is directly observed and essentially undisputed.
Whether the opt-out biases measured speedups downward turns on a premise that cannot be observed directly: that the developers who opt out would show larger AI speedups than those who participate. The case for it is that the selection appears driven by expected AI benefit rather than by the reduced pay rate, and that developers who refuse to work without AI have typically built their workflows around it, which tends to raise their measured speedup relative to unassisted work. The main reservation is that developers' self-assessments of AI-driven speedup are unreliable: in METR's own earlier study, participants believed AI had sped them up while it had in fact slowed them down, so a strong preference for working with AI does not by itself establish a larger true benefit.
On balance the direction of the bias is probable but not established. METR itself states the effect only as likely, and the surrounding commentary broadly accepts it, with some observers reframing the refusal to work without AI as dependency rather than evidence of higher productivity. Data from METR's redesigned experiment, or any measurement of opt-out developers under a design they will accept, would sharpen the picture.
Full reasoning: the evidence and decisions behind this verdict
The primary source is METR's February 2026 uplift update (metr.org/blog/2026-02-24-uplift-update/), which asserts the claim in its own voice: a significant increase in developers declining to participate because they do not wish to work without AI "likely biases downwards our estimate of AI-assisted speedup," and the true speedup "could be much higher among the developers and tasks which are selected out of the experiment." Both passages are recorded as affirming instances. Secondary commentary (philippdubach.com/posts/93-of-developers-use-ai-coding-tools.-productivity-hasnt-moved./) affirms the claim independently, describing the sample as selected toward developers who benefit least from AI; other discussion (Hacker News, commonplace blogs) reproduces METR's statement without contesting it. No published source was found that denies the claim outright.
The decomposition separates the observed premise from the counterfactual one. The observed premise, that AI-reliant developers increasingly decline to participate, rests on METR's direct recruitment experience and widely repeated refusal figures of 30 to 50 percent; it is effectively uncontested. The counterfactual premise, that opt-outs would show larger speedups than participants, is the crux: it cannot be tested on the opt-outs themselves. Two considerations support it: METR attributes the rising selection mainly to higher expected AI uplift rather than the pay cut, meaning attrition is selected on the treatment effect rather than on incidentals; and developers whose workflows are built around AI plausibly have both genuinely higher AI-assisted output and slower unassisted work, either of which raises their within-person measured speedup. Against it, self-reported speedup estimates are unreliable: METR's 2025 study found participants who were slowed roughly 19 percent believed they had been sped up about 20 percent, so expectation-driven opt-out is weak evidence of true uplift, and one commentator (paddo.dev) reframes the refusals as dependency rather than productivity. The miscalibration point weakens but does not defeat the mechanism, because the workflow-adaptation and skill-atrophy channels operate independently of belief accuracy.
Verdict: supported rather than verified. The behavioral premise is established, the direction-of-bias inference is the consensus reading and is more plausible than its alternatives, but it rests on an unobservable counterfactual that even the originating source hedges as "likely." Credence 0.8. What would change the conclusion: data from METR's redesigned experiment or any design that measures previously opting-out developers; evidence that opt-out correlates with belief but not with actual uplift (e.g., opt-outs induced to participate showing speedups no larger than participants') would push toward contested or contradicted.
Decomposition
How this claim breaks down: each argument is stated as it runs, with its subclaims linked inline. ↗︎ opens a subclaim; the map shows how they fit together.
Because AI-reliant developers increasingly decline to participate in experiments requiring AI-free work, and because those who opt out would experience larger AI speedups than those who participate, the sample that remains understates the average speedup, so measured estimates are biased downward. That the selection is driven mainly by expected AI uplift rather than by pay indicates the attrition is selected on the treatment effect itself rather than on incidental features of the study.
The inference is sound: if attrition is rising and is differential in the treatment effect, the remaining sample understates the average speedup. The observed premise that AI-reliant developers increasingly decline to participate is well documented; the argument's weight rests on whether the opt-outs would in fact show larger speedups, an unobservable counterfactual made plausible, though not established, by the finding that the selection is driven mainly by expected AI uplift rather than pay.
Because developers' self-reported estimates of AI-driven speedup are unreliable, developers who refuse to work without AI are not thereby shown to gain more from it, so the observed opt-out pattern need not produce a downward bias in the measured speedup.
Granting that developers' self-assessments of AI speedup are unreliable, opting out over a preference for AI is indeed weak evidence of higher true uplift, so the objection genuinely weakens the belief-based route to the bias. It does not reach the other routes: developers who refuse to work without AI have typically adapted their workflows around it and may work more slowly unassisted, both of which raise their measured speedup regardless of how accurate their beliefs are. The argument therefore reduces confidence in the size and certainty of the bias without showing the estimates are unbiased.
Provenance
Where this claim has been said, linked to its canonical form.
the true speedup could be much higher among the developers and tasks which are selected out of the experiment
Discussion after presenting raw quantitative results.
a significant increase in developers choosing not to participate in the study because they do not wish to work without AI, which likely biases downwards our estimate of AI-assisted speedup
Explanation of why the new experiment's signal is unreliable.
In February 2026, METR published an update changing their experiment design after discovering that 30-50% of invited developers declined to participate without AI access, a selection effect that biased the original sample toward developers who benefit least from AI.
A commentary on the gap between near-universal AI coding tool adoption and flat productivity statistics, arguing the exact magnitude of METR's July 2025 slowdown finding is now in doubt because the sample was selected toward developers who benefit least from AI.
Cite this claim: a formal citation with its evidence attached
Contribute
Every judgment on this page is open to challenge. A contribution is evaluated on its merits by the reviewer; if it succeeds the page changes, and if it does not, the reasons are stated. Either way the exchange becomes part of the claim’s public record.
The attention this claim received was paid for by a funded mandate. Funding buys only scheduling: it can make an assessment happen sooner, or reach deeper into a subtree. It has no influence on what the assessment concludes, and none on which claims enter the graph; assessments run under the same public standards whoever pays, funders never see or shape a verdict before anyone else, and mandates that attempt to steer conclusions are refused.
Created by extractor · Aug 10, 2026. Every judgment on this page is accompanied by a reasoning trace.