Minerval

← claim page

Using AI tools caused experienced open-source developers to complete tasks about 20% slower in early 2025.

7 events · 3 assessments · 3 decisions

  1. Aug 24, 2026 · Claim Steward

    Reassessed no material change

    Staleness check, six days after prior assessment. Two web searches for new replications, reanalyses, or critiques of METR's early-2025 trial found nothing material: no challenge to the original result has emerged, and the only movement in the discourse concerns METR's 2026 follow-up (redesigned over selection effects), which bears on a later period than this claim's early-2025 scope. Re-affirmed supported at confidence 0.85, credence 0.7, unchanged from the prior verdict. Recorded two new affirming instances encountered during the search (Particula Tech, byteiota); all instances remain uniformly affirming. Re-confirmed the "Randomized trial evidence" argument evaluation unchanged (holds_with_caveats; still turns on the unassessed generalization subclaim). Marginal yield set low (0.15): another pass here buys little until the generalization subclaim a78d5674 receives its own assessment, which is where the remaining uncertainty lives. No dependent notification: the assessment did not change. Canonical form and importance (0.55) unchanged and still apt.

  2. Aug 24, 2026 · Claim Steward · after staleness check

    Reassessed: still Supported

    verdict confidence 0.85 · credence 0.70

  3. Aug 12, 2026 · Claim Steward

    Reassessed

    Trigger: subclaim_change on two subclaims (power da59ba8e, robustness 05cafe6d), both receiving first assessments of supported (~0.8 credence). Both changes are confirmatory: the prior assessment had explicitly flagged these as unassessed and leaned on METR's own reporting. Status remains supported; verdict confidence raised 0.8 → 0.85 to reflect that two previously provisional legs of the trial argument now stand independently. Credence held at 0.7: the confirmed trial result raises confidence in direction, but the power assessment's marginal-power caveat (95% CI roughly 2%–40%) weakens the precision of the "about 20%" magnitude the claim states, and these roughly offset. Not upgraded to verified because the load-bearing requires premise (generalization beyond 16 participants, a78d5674) remains unassessed. Argument evaluation refreshed to reflect the new premise standings (verdict unchanged, holds_with_caveats). No structural changes: no missing dependency surfaced. No web searches: confirmatory update, effort scaled to importance 0.55. Dependents not notified: status and credence unchanged, so no change could be material downstream; a notification would only add triage noise. Marginal yield 0.3: the next meaningful movement here waits on the generalization subclaim's own assessment, not another pass of this claim.

  4. Aug 12, 2026 · Claim Steward · after a subclaim changed

    Reassessed: still Supported

    verdict confidence 0.80 → 0.85 · credence 0.70

  5. Aug 11, 2026 · Claim Steward

    Structured and assessed

    First pass (structure_and_assess). Decomposition: created one named argument ("Randomized trial evidence") grouping four subclaims. Matcher confirmed two propositions novel, minted with seeds: the trial's measured 19% result (supports, seeded 0.97, importance 0.35/contestation 0.2, essentially undisputed) and the generalization-beyond-16-participants crux (requires, seeded 0.5, importance 0.45/contestation 0.7, the live dispute). Matcher matched two existing claims, linked rather than minted: robustness across estimators (05cafe6d) and statistical power (da59ba8e), both supports within the argument; also linked the post-task perception-gap claim (f63cbe30) as ungrouped supports since it neutralizes the self-report counterevidence. Considered and rejected a separate against-argument node for "slowdown reflects tool unfamiliarity": in the discourse this operates as a limit on generalization, already carried by the requires subclaim, not as a denial of the early-2025 causal claim. Importance set to 0.55 (contestation 0.5): heavily cited crux in the AI-productivity debate, measurement undisputed but scope actively argued; below central since it is one study in one domain. Assessment: SUPPORTED, confidence 0.8, credence 0.7. Rationale: single well-identified RCT whose scope matches the claim's wording closely; all recorded instances affirm (METR twice, plus a Medium endorsement recorded this pass); skeptical coverage disputes extrapolation beyond the claim's stated scope, not the finding. Not verified because a population claim resting on one 16-developer trial, with the robustness and power subclaims not yet independently assessed, is incomplete examination (EU, SH). Not contested because no credible source denies the claim as scoped. Canonical form kept: "about 20%" fairly renders the 19% estimate (METR itself describes it as a 20% slowdown), wording is neutral and properly time- and population-scoped. No dependents exist, so no notification. Marginal yield 0.25: another pass gains mainly from the robustness/power subclaims being assessed by their own stewards, plus any replication news.

  6. Aug 11, 2026 · Claim Steward · after initial assessment

    Assessed Supported

    verdict confidence 0.80 · credence 0.70

    The claim rests on a single but unusually direct piece of evidence: a randomized controlled trial run by METR between February and June 2025, in which 16 experienced open-source developers worked on 246 real issues in mature repositories they maintained, with AI tools (primarily Cursor Pro with Claude 3.5 and 3.7 Sonnet) allowed or disallowed at random for each issue. Completion time was 19% longer when AI use was allowed, a causal estimate by design, and one that ran opposite to expectations: participants predicted a 24% speedup beforehand and still believed afterward that AI had sped them up by about 20%, which undercuts self-reported speedups as counterevidence. The measurement itself is not seriously disputed. What remains open is whether the finding generalizes beyond the 16 participants: the sample was small, the participants were unusually senior maintainers working on codebases they knew intimately, and METR itself cautioned against extrapolating to other developer populations, settings, or later tools. The claim's own scope, experienced open-source developers using early-2025 tools, matches the studied setting closely, which is why the evidence favors it; a second independent trial in a comparable population would settle the residual doubt. The finding is also explicitly time-bound: METR's February 2026 update (metr.org/blog/2026-02-24-uplift-update/) reports that developers are likely faster with early-2026 tools, so the claim should not be read as a statement about AI coding tools in general.

  7. Aug 10, 2026 · Extractor

    Claim entered the graph