METR's 2025 sample of 246 completed issues from 16 developers gave sufficient statistical power to detect AI's effect on developer productivity.
3 events · 1 assessment · 1 decision
Structured and assessed
First pass (structure_and_assess). Decomposition: minted three subclaims after match_claim confirmed novelty — the effective-sample premise (requires, e63f5522, the crux where the "only 16 developers" objection lives), significance under developer-level clustering (supports, 6fc6cd00), and the detection-versus-magnitude specification (specifies, bac0ee12) — and linked the existing robustness claim 05cafe6d (supports) that the Matcher surfaced. One natural line of support, so no named argument. Evidence: METR's methodology FAQ, the Kucinskas independent reanalysis (clustered SEs + hierarchical bootstrap, findings unchanged), and a participant's report of the 95% CI (≈2%–40% slowdown). Recorded the Kucinskas post as an affirming instance here (conf 0.7) and on the clustering subclaim (conf 0.85). Verdict: supported, confidence 0.8, credence 0.8; not verified because the power argument is post-hoc, the CI excludes zero only narrowly, and the low within-developer-correlation premise rests on METR's report plus one reanalysis. Marginal yield 0.2: reading Appendix D and the reanalysis notebook directly could sharpen but likely not change the verdict. Canonical form updated to name METR (was "The study's...", ambiguous out of context; same proposition). Importance set 0.4 (contestation 0.45): notable methodological premise in a live debate, largely settled technically but widely repeated in the discourse. Per V/SH: primary source (METR) checked against independent reanalysis; per ND: the against-consideration (marginal power, wide CI) is represented as its own subclaim.
Assessed Supported
verdict confidence 0.80 · credence 0.80
The claim concerns the most common objection to METR's 2025 finding that AI tools slowed experienced open-source developers: that a study of only 16 people cannot support statistical conclusions. The evidence favors the claim, with an important qualification about what "sufficient power" covered. The study randomized AI use at the level of individual issues, with every developer working in both conditions, so the effective sample for the average effect was the 246 completed issues rather than the 16 developers, provided within-developer correlation is modest, which METR reports it is. The decisive test is whether the result survives inference that respects the developer clusters, and it does: the measured slowdown remained statistically significant under developer-clustered standard errors, a result confirmed by an independent reanalysis of the open-sourced data that also applied a hierarchical bootstrap, the method of choice when clusters are few. More broadly, the finding is robust across alternative estimators and empirical strategies. The qualification is that the power was marginal and specific to detection: the 95% confidence interval ran from roughly a 2% to a 40% slowdown, so the study could establish the existence and direction of the effect but not its magnitude with any precision, and per-developer or subgroup estimates are noisier still. METR's own characterization, "just enough" power to reject zero effect, is accurate. Readers who cite the study for the existence of a slowdown in this setting are on solid statistical ground; readers who lean on the 19% point estimate as a precise quantity are not. Whether the result generalizes beyond these developers and projects is a separate question from power and is treated as its own claim.
Claim entered the graph