AI tools sped up developers more in early 2026 than early-2025 estimates indicated.
Assessment
Evidence favors the claim, but the chain is incomplete or the sources are secondary.
The claim originates with METR, the research group whose early-2025 randomized study found that AI tools made experienced open-source developers about 19% slower. In its February 2026 update, the same team reported follow-up data with point estimates on the speedup side, roughly an 18% speedup for returning developers and 4% for new recruits, and stated it is likely that developers are more sped up by AI tools in early 2026 than the early-2025 estimates indicated.
The evidence favors the claim but falls short of establishing it. The follow-up experiment yields an unreliable signal: confidence intervals span zero, participation dropped, and developers withheld tasks they did not want to do without AI. Yet those selection effects bias the measured speedup downward, since the developers and tasks with the highest expected uplift are the ones missing, so the flawed measurements function as a conservative floor rather than a refutation. Substantial capability gains in agentic coding tools during 2025 add outside plausibility. For the claim to be false, the true early-2026 effect would have to sit at or below the early-2025 finding of a roughly 20% slowdown, a reading no credible source asserts.
A reliable direct measurement is what would settle the question. METR has redesigned its experiment in response to the selection problems; results from that redesign, or an independent randomized study, would either confirm the increase or contradict it.
Full reasoning: the evidence and decisions behind this verdict
Staleness re-check, five days after the prior assessment. Searches for new METR publications and independent measurements found nothing that moves the verdict: the most recent primary source remains METR's February 2026 update (metr.org/blog/2026-02-24-uplift-update/), and commentary published since the last pass (e.g. valueaddvc.com/blog/ai-coding-productivity-study-data-what-metr-mckinsey-and-github-actually-found-in-2026, July 2026; scienceblog.com/, July 2026) reports METR's existing estimates rather than adding new data. METR's redesigned experiment has not published results. Neither of the conditions the prior assessment named as verdict-changing has occurred: no redesigned-experiment or independent randomized result at or below the early-2025 estimate (which would contradict the claim), and no reliable measurement above it (which would move it toward verified).
The evidential picture is therefore unchanged. The follow-up data's raw numbers (-18%, CI -38% to +9%, returning developers; -4%, CI -15% to +9%, new recruits) sit above the early-2025 estimate (+19%, CI +2% to +39%) but carry intervals spanning zero, and the follow-up experiment's unreliability (supported, 0.85) blocks treating them as a measurement. The directional case rests on the downward selection bias from AI-reliant developers opting out: 30-50% of surveyed developers withheld tasks, participation without AI declined, and METR states the new estimates are likely a lower bound. Tool capability gains during 2025 supply prior plausibility; the early-2025 slowdown finding (supported, 0.85) fixes the comparison baseline. All recorded instances affirm; no credible source asserts the negation.
Status stays supported rather than verified because every quantitative estimate involved has a confidence interval spanning zero and the originating team itself hedges ("likely"); it stays supported rather than contested because measurements known to be biased downward already sit on the speedup side and no credible dissent exists. Credence 0.85, as before: the claim inherits the residual doubt of its small-sample baseline, and the case is a bias argument plus one team's qualitative judgment, not a measurement. Marginal yield is low: another pass buys little until METR's redesigned experiment or an independent study publishes; that publication, not more analysis of existing data, is what would change this assessment. METR's May 2026 self-reported usage survey (metr.org/blog/2026-05-11-ai-usage-survey/) remains a candidate source, weighed against the documented unreliability of self-reported speedups.
Decomposition
How this claim breaks down: each argument is stated as it runs, with its subclaims linked inline. ↗︎ opens a subclaim; the map shows how they fit together.
The claims this one rests on directly, not gathered into a named line of reasoning.
- assumesbackground the parent's framing takes as givensteward instructions →Using AI tools caused experienced open-source developers to complete tasks about 20% slower in early 2025. ↗︎
- supportsthis provides evidence for the parentsteward instructions →AI coding tools became substantially more capable during 2025 with the adoption of agentic tools. ↗︎
METR's follow-up experiment produced raw estimates of roughly an 18% speedup among returning developers and a 4% speedup among newly recruited ones, both already above the early-2025 finding of a roughly 19% slowdown. Because selection effects from AI-reliant developers opting out bias these measured estimates downward, the true early-2026 effect likely sits above even those raw figures, and therefore above the early-2025 estimates.
The inference is sound: if the measured estimates are biased downward and already sit above the early-2025 slowdown, the true effect sits above the early-2025 estimates. Its weight rests almost entirely on the downward selection bias from AI-reliant developers opting out, which is not yet independently assessed but rests on a straightforward mechanism the experimenters describe directly: the developers and tasks with the highest expected uplift are the ones systematically missing. The residual weakness is that the raw estimates' confidence intervals span zero, so the argument establishes direction more firmly than magnitude.
Because the follow-up experiment yields an unreliable signal of AI's current productivity effect, and its raw estimates carry confidence intervals spanning zero, the new data cannot by itself establish that developers are more sped up in early 2026, leaving the claim to rest on the experimenters' qualitative judgment and indirect considerations.
The premise stands well supported and the inference goes through as far as it reaches: the follow-up experiment's unreliability genuinely blocks any conclusion drawn from the raw speedup estimates alone, which is why the claim cannot be assessed as verified. The caveat is that the argument weighs against certainty rather than against the claim's direction: the dominant source of the unreliability is a selection bias known to push the estimates downward, so the same defect that disqualifies the data as a measurement leaves it usable as a conservative floor.
Provenance
Where this claim has been said, linked to its canonical form.
we believe it is likely that developers are more sped up from AI tools now — in early 2026 — compared to our estimates from early 2025
Discussion following description of selection effects and pay changes.
But the evidence is moving from "net negative" toward "modest gains, not transformative."
A post arguing the "10x AI developer" is a myth, caveating its reliance on METR's early-2025 slowdown finding by acknowledging METR's February 2026 update: the updated cohort's estimates and METR's own statement that AI likely provides productivity benefits in early 2026.
The honest takeaway is that 2026 is a transitional year: agentic AI is pulling the productivity curve upward, but only for developers who have already diagnosed and fixed the context-switching and measurement problems underneath.
A commentary on METR's February 2026 update that reports METR's belief that developers are more sped up in early 2026 than early-2025 estimates indicated, credits agentic tools (Claude Code, Codex) for the shift, and endorses the upward direction in its own voice with workflow caveats.
Assessment history
0 status changes over 2 assessments. full history →
Cite this claim: a formal citation with its evidence attached
Contribute
Every judgment on this page is open to challenge. A contribution is evaluated on its merits by the reviewer; if it succeeds the page changes, and if it does not, the reasons are stated. Either way the exchange becomes part of the claim’s public record.
The attention this claim received was paid for by a funded mandate. Funding buys only scheduling: it can make an assessment happen sooner, or reach deeper into a subtree. It has no influence on what the assessment concludes, and none on which claims enter the graph; assessments run under the same public standards whoever pays, funders never see or shape a verdict before anyone else, and mandates that attempt to steer conclusions are refused.
Created by extractor · Aug 10, 2026. Every judgment on this page is accompanied by a reasoning trace.