AI coding assistants substantially speed up developers on many programming tasks.
3 events · 1 assessment · 1 decision
Structured and assessed
First pass (structure_and_assess). Decomposition: created two named arguments. For: "Controlled evidence of speedups" holding two newly minted subclaims after match_claim returned novel for both (GitHub 2022 experiment ~55% faster, seeded 0.9; Copilot field experiments ~26% more tasks, seeded 0.8), plus two existing claims linked via add_relationship_edge (heterogeneity of gains 75326d90, early-2026 speedup vs early-2025 estimates aca2b03a). Against: "Slowdown in experienced, familiar-codebase settings" holding the existing METR slowdown claim 7cbe87e8 (contradicts). Both arguments given written forms and evaluations (both holds_with_caveats). New subclaims scored 0.3-0.4 importance: specific study-result claims, moderately contested (interpretation more than measurement). Evidence pass: four web searches confirmed Peng et al. arXiv:2302.06590 (55.8% faster, CI 21-89%), Cui et al. field experiments (~26% more tasks, PR-based, strongest at Microsoft), ANZ trial (42% faster), METR RCT arXiv:2507.09089 (19% slowdown for experienced OSS maintainers), and METR's Feb 2026 update (follow-up data unreliable due to selection effects; team believes speedup likely larger now). One affirming instance recorded (a literature-review preprint endorsing the productivity conclusion in its own voice) at 0.6 confidence given the productivity-vs-speed phrasing gap. Verdict: SUPPORTED, confidence 0.7, credence 0.75. Contested was the serious alternative (credible RCTs point in opposite directions) but was set aside because the disagreement targets a stronger claim (speedup on most/all work including expert work in mature codebases); as worded, "many programming tasks" is covered by the favorable evidence and bounded, not negated, by METR. Importance set to 0.6 (was 0.55), contestation 0.7: major-tier, live debate, feeds AI-capability discourse. Marginal yield 0.35: 2026-era evidence is still unfolding and a future pass with newer RCT data could sharpen or flip the scope judgment. Canonical form kept: 12 words, neutral, acceptable to both sides.
Assessed Supported
verdict confidence 0.70 · credence 0.75
Controlled experiments have repeatedly found large speedups from AI coding assistants in a wide range of settings. GitHub's 2022 randomized experiment found Copilot users finished a self-contained coding task about 55% faster, and large randomized field deployments at Microsoft, Accenture, and a third firm found roughly 26% more tasks completed in ordinary enterprise work, with a six-week trial at ANZ Bank reporting completion roughly 42% faster. Gains in these studies were largest for less experienced developers and for work in unfamiliar territory, conditions common in real software work. The main counterevidence is a 2025 randomized trial by METR in which experienced open-source maintainers working on large codebases they knew well completed tasks about 20% slower with AI tools, even while believing the tools had sped them up. That result shows the speedup is not universal: it can vanish or reverse for expert work in mature, familiar codebases. It bounds the claim rather than refuting it, since the favorable evidence covers greenfield tasks, enterprise development, and less experienced developers, which plausibly constitute many programming tasks. There is also indication that tools in early 2026 sped developers up more than early-2025 measurements suggested, though the follow-up data behind that indication suffered from selection problems. The credible disagreement is therefore about scope, not existence: how representative the settings with large measured gains are of real-world programming as a whole, and whether measured speedups on individual tasks translate into overall productivity. Continued randomized measurement across task types and experience levels, of the kind METR has attempted to run, would sharpen the answer.
Claim entered the graph