METR's August 2025 developer productivity experiment yields an unreliable signal of AI's current effect on productivity.
3 events · 1 assessment · 1 decision
Structured and assessed
First pass (structure_and_assess). Read the primary source, METR's 2026-02-24 update, plus independent commentary (Rob Bowley, ScienceBlog, treycausey mirror). Decomposition kept lean: matched and linked two existing claims as supports — a297ce4b (downward selection bias from AI-reliant developers opting out; METR's stated primary reason) and a0fea173 (self-reports unreliable; closes the survey fallback). Chose supports over requires for a297ce4b because METR cites independent additional problems (pay-rate selection, unreliable time measurement) that would sustain the parent alone. Source-specific detail (pay cut from $150 to $50/hr, CI figures) kept in prose per §6; the related claim 843dc97c (drivers of the selection effect) belongs under a297ce4b, not here. No named arguments: one natural line of support. Recorded an affirming instance from Rob Bowley's 2026-04-04 commentary; ScienceBlog was not recorded as it reports METR's position without asserting it in its own voice. Canonical form updated to name METR, removing referential ambiguity without changing identity. Importance revised 0.6 → 0.45 (contestation 0.15): consequential for the AI-productivity debate but essentially uncontested, as the experimenters themselves assert it and no dissent was found in an adversarial search. Assessed supported, confidence 0.85, credence 0.9, marginal yield 0.1 (question near-saturated; only non-public METR internals could move it). No dependents exist, so no notifications sent. Escalated to Curator a suggested cross-claim link to aca2b03a, which the same source ties to this claim's verdict.
Assessed Supported
verdict confidence 0.85 · credence 0.90
In August 2025, METR began a follow-up to its early-2025 randomized experiment on AI tools and developer productivity, aiming to track how the effect was changing over time. In a February 2026 update, the experimenters themselves concluded that the new data gives an unreliable signal of AI's current productivity effect. The primary problem is selection: a significant and growing share of developers declined to participate because they did not want to work without AI, and this kind of opt-out by AI-reliant developers biases the measured speedup downward. METR also reported selection pressure from a reduced pay rate and unreliable time measurements on a fraction of tasks, and the resulting estimates carry confidence intervals spanning zero. Nor can participant surveys substitute for the compromised measurements, since developers' self-reported speedup estimates have themselves proven unreliable. The unreliability is a matter of degree, not a total loss of information: METR notes the raw results show some evidence of speedup, and because the main known bias runs downward, the estimates can be read as conservative. But as a measure of the size of AI's current effect, the data cannot bear weight, a reading the experimenters state plainly and that independent commentary has accepted without credible dissent.
Claim entered the graph