AI tools sped up developers more in early 2026 than early-2025 estimates indicated.
5 events · 2 assessments · 2 decisions
Reassessed no material change
Staleness check five days after the prior assessment. Two web searches confirmed the evidence landscape has not moved: METR's February 2026 update remains the most recent primary source, its redesigned experiment has not published results, and no independent randomized measurement has appeared. Post-August commentary found (valueaddvc.com, scienceblog.com, treycausey.com) reports METR's existing estimates rather than asserting new findings in its own voice, so no new instances were recorded. Neither verdict-changing condition named in the prior assessment has occurred. Re-affirmed supported at confidence 0.8, credence 0.85, with both argument evaluations re-recorded unchanged. Marginal yield set low (0.15): further passes buy little until METR's redesign or an independent study publishes. Structure, canonical form, and importance (0.6) all left unchanged; no notification to dependents since the assessment did not change.
Reassessed: still Supported
verdict confidence 0.80 · credence 0.85
Structured and assessed first pass
First pass (structure_and_assess) with Curator edge suggestions. Decomposition: adopted all three suggested claims with the following renderings. (1) 7cbe87e8 (early-2025 ~20% slowdown) attached as ASSUMES at claim level: it is the comparison baseline; if it failed the claim would be mischaracterized rather than false. (2) a297ce4b (downward selection bias) attached as SUPPORTS under a new FOR argument "Direction-of-bias reading of the follow-up data": it is the load-bearing premise converting unreliable raw data into directional evidence. (3) 10c1391b (unreliable signal) attached as CONTRADICTS under a new AGAINST argument "Unreliability of the direct measurements": the Curator noted it cuts both ways; I rendered the undercutting side as the against-argument (evaluated holds_with_caveats, since it blocks magnitude, not direction) and let the bias-direction side live in the for-argument. Minted one new subclaim after match_claim returned novel (confidence 0.92): "AI coding tools became substantially more capable during 2025 with the adoption of agentic tools" (supports, importance 0.3, contestation 0.35, seed 0.85), giving the claim prior plausibility independent of METR's compromised experiment. Evidence: read METR's Feb 2026 update directly (raw estimates -18%/-4% change in completion time, CIs spanning zero; opt-out, pay-cut, and task-withholding selection pressures; METR's own "likely" conclusion), plus commentary (paddo.dev, PorkiCoder, Rob Bowley) and a LessWrong reanalysis estimating a sample-wide 6% speedup. Recorded two affirming instances (paddo.dev, PorkiCoder) encountered during evidence reading; skipped the substack cross-post (duplicate of the metr.org source) and Rob Bowley (reports the study breakdown without asserting the claim). Assessment: SUPPORTED, confidence 0.8, credence 0.85, marginal_yield 0.3 (METR's May 2026 usage survey and any redesigned-experiment results are undigested future evidence). Not verified: all estimates have zero-spanning CIs and the originating team hedges. Not contested: no credible source asserts the negation, and falsity would require the true early-2026 effect to sit below biased-downward measurements' point estimates. Also updated canonical form ("AI tools sped up developers more in early 2026 than early-2025 estimates indicated", fixing the awkward "than according to" phrasing, same proposition) and set importance 0.6 / contestation 0.45, near the Extractor's 0.65 prior. Notifying the one dependent (89098aaf, benchmark-overestimation claim, linked by contradicts) since this is the claim's first assessment.
Assessed Supported
verdict confidence 0.80 · credence 0.85
The claim originates with METR, whose early-2025 randomized experiment found that AI tools slowed experienced open-source developers by about 20%. In a February 2026 update (metr.org/blog/2026-02-24-uplift-update/), the same team reported that its follow-up experiment, begun in August 2025 with newer tools, produced raw estimates of roughly an 18% speedup among returning developers and a 4% speedup among newly recruited ones, both with confidence intervals spanning zero, and concluded that developers are likely more sped up in early 2026 than the early-2025 estimates indicated. The case for the claim is indirect but coherent. METR itself judges that the follow-up experiment yields an unreliable signal of AI's current productivity effect, so the raw estimates cannot carry the claim on their own. But the principal source of that unreliability points in a known direction: developers who rely on AI opting out of the study biases the measured speedup downward, which makes the raw figures a plausible lower bound, and even those figures sit well above the early-2025 slowdown finding. The shift is also what the tool landscape would predict, since AI coding tools became substantially more capable during 2025 with the adoption of agentic tools. Commentary on the update has accepted the directional conclusion, and no credible source asserts that developers were equally or less sped up in early 2026. What keeps the claim short of established is that no reliable measurement of the early-2026 effect yet exists: the conclusion rests on the experimenters' own hedged judgment, the direction-of-bias reasoning, and weak raw data. METR is redesigning its experiment in response; a completed redesigned experiment, or an independent randomized study of comparable developers with early-2026 tools, would settle the size of the change and harden or overturn the direction.
Claim entered the graph