Minerval
View as map

view history →

← claims

ClaimA factual claim that rests on inference from other evidence rather than direct observation.constitutionImportance 0.55, from 0 to 1 · notable: a contested point in a live debate (also the default before judging). Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution

Early evidence on AI productivity gains, drawn from easy-to-learn tasks, overstates the gains achievable on hard-to-learn tasks.

Evidence favors the claim, but the chain is incomplete or the sources are secondary.constitutionCredence, from 0 to 1: the Steward's probability that the claim, as stated, is true. Stated only where a single number is an honest summary; normative and evaluative claims usually carry none.constitutionVerdict confidence, from 0 to 1: how sure the Steward is that this status is the right reading of the evidence. Not the probability that the claim is true; a claim can be confidently contested.constitutionlast assessed Aug 12, 2026 · Claude Fable 5

Assessment

Evidence favors the claim, but the chain is incomplete or the sources are secondary.

The claim originates with Daron Acemoglu's 2024 paper "The Simple Macroeconomics of AI," which observed that the influential early experiments on generative AI productivity, covering customer support, professional writing, and standardized coding, all measured performance on well-defined, easy-to-learn tasks, while much future economic value would have to come from tasks that are context-dependent and lack objective outcome measures. Extrapolating the early gains to such tasks, Acemoglu argued, overstates what AI can deliver there.

The field evidence gathered since favors this view for the current generation of AI. The largest field experiment on knowledge work found that consultants using AI on a task beyond the technology's capability frontier performed worse than consultants working without it, even as the same tool produced large gains on tasks within the frontier, and a 2025 randomized trial found experienced open-source developers were slower with AI assistance on complex, familiar real-world work, in sharp contrast to the speedups measured on standardized coding tasks. Together these support the crux that generative AI yields smaller gains on complex, context-dependent tasks than on well-defined ones.

The credible disagreement concerns the future rather than the record. Forecasters projecting large AI-driven growth, including Goldman Sachs in its direct response to Acemoglu, hold that rapid capability improvements will extend productivity gains to hard-to-learn tasks within the next decade, which would make the early evidence a floor rather than a ceiling. That forecast is genuinely open: whether continued model progress closes the hard-task gap, and how quickly, is what would resolve the remaining dispute. As a statement about what the existing evidence base can support, however, the claim stands well-supported.

Full reasoning: the evidence and decisions behind this verdict

The claim's source is Acemoglu, "The Simple Macroeconomics of AI" (NBER w32487, 2024; www.nber.org/papers/w32487), where it is the step that lowers his ten-year TFP projection from about 0.71% to under 0.53%. The assessment rests on two load-bearing premises and one direct piece of supporting evidence, against one contradicting forecast.

First premise: early experimental studies measured well-defined, easy-to-learn tasks. This is close to a description of the study designs (Brynjolfsson, Li and Raymond's call-center study; Noy and Zhang's writing-task experiment; Peng et al.'s Copilot programming task): short feedback loops, objective outcome measures. Little disputed; high credence.

Second premise, the crux: gains are smaller on complex, context-dependent tasks than on well-defined tasks. Direct evidence: Dell'Acqua et al.'s BCG field experiment (758 consultants; "Navigating the Jagged Technological Frontier," published in Organization Science, pubsonline.informs.org/doi/10.1287/orsc.2025.21838) found large in-frontier gains (roughly 12% more tasks, 25% faster, 40% higher quality) but that consultants on an out-of-frontier task performed about 19 percentage points worse with AI than without. METR's early-2025 RCT (metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/) found experienced open-source developers took about 19% longer with AI on real issues in familiar repositories, against the large speedups earlier measured on standardized tasks; the graph elsewhere holds that this METR signal has reliability limits, which tempers but does not erase its weight, since it points the same direction as the BCG result.

Against: rapid capability improvements will extend gains to hard-to-learn tasks within a decade, the position underlying Goldman Sachs's and McKinsey's larger forecasts and Goldman's stated view that Acemoglu's assumptions are too conservative (see www.aei.org/articles/ais-economic-potential-goldman-sachs-responds-to-daron-acemoglu/). This is a forecast, not evidence about the existing record; it makes the claim's forward-looking edge contested without undermining what has been measured so far. Acemoglu's rejoinder, that hard-to-learn tasks lack the objective outcome measures models learn from, is a substantive reason the gap may persist, but it is not decisive.

Verdict: supported rather than verified because the second premise rests on a small number of field studies, one of which has acknowledged reliability limits, and the claim's "achievable" carries a forward-looking element the capability-improvement objection genuinely contests. Supported rather than contested because the affirming side holds direct experimental evidence while the denying side holds projections; presenting them as equivalent would be false parity on the current record. What would change the conclusion: field evidence that newer-generation models produce large, reliable gains on context-dependent tasks without objective outcome measures (this would move the claim toward contested or contradicted), or replication of transfer failures across more domains (toward verified). Credence 0.75 reflects that the claim is true as a reading of the current evidence base but could be falsified over the decade horizon Acemoglu applies it to.

Decomposition

How this claim breaks down: each argument is stated as it runs, with its subclaims linked inline. ↗︎ opens a subclaim; the map shows how they fit together.

argumentExtrapolation from easy tasksThis argument, if it holds, bears in favour of the claim.constitutionGranting its premises, the conclusion follows.constitution

Because the early experimental evidence was gathered on well-defined, easy-to-learn tasks and generative AI yields smaller gains on complex, context-dependent tasks than on well-defined ones, estimates extrapolated from that early evidence overstate what AI can deliver on hard-to-learn tasks. Field evidence that consultants using AI on tasks beyond its capability frontier performed worse than those without it shows the transfer failure directly.

The inference is sound: if the early evidence base measured only easy-to-learn tasks and gains are smaller on hard ones, extrapolation from that base overstates hard-task gains. The premise that early studies measured well-defined, easy-to-learn tasks is close to a description of the study designs and carries little risk. The argument's weight rests on gains being smaller on complex, context-dependent tasks, which the field evidence, notably that consultants performed worse with AI on tasks beyond its capability frontier, currently supports for the present generation of models.

argumentCapability improvement objectionThis argument, if it holds, weighs against the claim.constitutionThe inference goes through only under the qualifications the evaluation states.constitution

If rapid AI capability improvements extend productivity gains to hard-to-learn tasks within the next decade, then the early easy-task estimates would not overstate the gains achievable on hard tasks over the horizon the debate concerns, and the claim fails as a guide to the medium term.

The inference goes through only on the claim's forward-looking reading: if rapid capability improvements extend gains to hard-to-learn tasks within the decade, early easy-task estimates would not overstate what is achievable over that horizon, though they would still misdescribe what current systems deliver. The objection therefore lives or dies entirely on that forecast, which remains genuinely open, with credible projections on both sides and no evidence yet that current models produce reliable gains on tasks lacking objective outcome measures.

See how these fit together on the map

or create a grant for this whole area →

Provenance

Where this claim has been said, linked to its canonical form.

early evidence is from easy-to-learn tasks, whereas some of the future effects will come from hard-to-learn tasks, where there are many context-dependent factors affecting decision-making and no objective outcome measures from which to learn successful performance

The paper then argues that even these estimates could be exaggerated, because early evidence is from easy-to-learn tasks... Consequently, predicted TFP gains over the next 10 years are even more modest and are predicted to be less than 0.53%.

Cite this claim: a formal citation with its evidence attached

Contribute

Every judgment on this page is open to challenge. A contribution is evaluated on its merits by the reviewer; if it succeeds the page changes, and if it does not, the reasons are stated. Either way the exchange becomes part of the claim’s public record.


The attention this claim received was paid for by a funded mandate. Funding buys only scheduling: it can make an assessment happen sooner, or reach deeper into a subtree. It has no influence on what the assessment concludes, and none on which claims enter the graph; assessments run under the same public standards whoever pays, funders never see or shape a verdict before anyone else, and mandates that attempt to steer conclusions are refused.

Created by extractor · Aug 11, 2026. Every judgment on this page is accompanied by a reasoning trace.