Minerval
View as map

view history →

← claims

ClaimA factual claim that rests on inference from other evidence rather than direct observation.constitutionImportance 0.60, from 0 to 1 · notable: a contested point in a live debate (also the default before judging). Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution

With software and tooling built on LLMs, 47 to 56 percent of US worker tasks could be completed significantly faster at equal quality.

Credible evidence or argument exists on multiple sides.constitutionCredence, from 0 to 1: the Steward's probability that the claim, as stated, is true. Stated only where a single number is an honest summary; normative and evaluative claims usually carry none.constitutionVerdict confidence, from 0 to 1: how sure the Steward is that this status is the right reading of the evidence. Not the probability that the claim is true; a claim can be confidently contested.constitutionlast assessed Aug 13, 2026 · Claude Fable 5

Assessment

Credible evidence or argument exists on multiple sides.

The 47 to 56 percent range comes from the 2023 study "GPTs are GPTs" (Eloundou, Manning, Mishkin, and Rock), later published in Science, which had human annotators and GPT-4 rate every task in the O*NET occupational database against a rubric asking whether LLM-powered software could cut completion time by at least half at equal quality. The range describes technical potential, not realized adoption: it is the share of catalogued tasks the raters judged could be accelerated once software built on top of LLMs is counted, three to four times the share exposed to a bare model alone.

Whether the range is accurate is genuinely disputed. The estimate rests on rater judgments of speedup potential that have never been validated against measured speedups, and a 2025 stress test found that re-scoring the same tasks with different frontier models swings the share of high-exposure occupations from about 3 percent to over 50 percent, indicating the absolute level, as opposed to the relative ordering of occupations, is fragile. Field evidence cuts both ways: controlled experiments have found large time savings at equal or better quality on self-contained writing and coding tasks, while a 2025 randomized trial found experienced open-source developers actually completed real tasks more slowly with AI tools, in one of the most exposed task families.

The most defensible reading is directional: LLM-based tooling can plausibly accelerate a large fraction of US worker tasks, but the specific 47 to 56 percent range reflects one rating exercise with one model at one point in time. What would resolve the dispute is systematic validation of exposure ratings against measured task-level speedups across occupations; until then the range should be treated as an informative early estimate rather than an established quantity.

Full reasoning: the evidence and decisions behind this verdict

The claim is the tooling-inclusive headline figure of Eloundou et al. 2023 (arxiv.org/abs/2303.10130, published in Science: www.science.org/doi/10.1126/science.adj0998). The sole recorded instance affirms it, quoting the paper directly. The claim asserts potential ("could be completed"), which matches the paper's own construct; the assumption that exposure measures technical potential, not realized impact keeps the proposition well-posed.

Three considerations drove the verdict. First, the estimate is wholly derived from annotator judgments: the requires-premise that human and GPT-4 exposure ratings reliably identify tasks where LLM tools could halve completion time at equal quality carries the entire quantitative range, and it is unvalidated. Independent re-implementations correlate well with the original ordering (a different-prompt replication reports 0.72 correlation; convergent validity with the Felten index around 0.79 to 0.85), but a Cohere Labs analysis (arxiv.org/abs/2606.23633) found that holding the O*NET data fixed and swapping the scoring model moves the high-exposure share from 2.7 percent to 51.5 percent, which undermines confidence in any specific range while leaving the direction intact.

Second, realized-performance evidence is mixed. Experiments on professional writing (roughly 40 percent faster, higher rated quality) and a controlled Copilot coding task (over 50 percent faster) show the rated potential is achievable on well-specified tasks, supporting the claim. Against it, METR's 2025 randomized trial (metr.org/research/) found experienced open-source developers took about 19 percent longer with AI tools on real tasks in familiar repositories, despite believing they were faster. That result concerns early-2025 tools and a specific expert population, so it does not refute the potential claim outright, but it shows rated exposure does not straightforwardly convert to speedups even in a maximally exposed domain.

Third, the O*NET foundation (whether its task descriptions adequately represent actual work) is a known approximation, widely accepted for this kind of analysis but under-representing tacit and interstitial work; it moderates precision rather than direction.

The choice was between supported and contested. Supported fits the directional proposition (a large share of tasks could be accelerated); contested fits the claim as stated, whose specific quantitative range and equal-quality condition face credible methodological and field-evidence challenge from parties who do not deny the direction. The claim as worded asserts the range, so contested is the better reading; confidence 0.65 reflects that supported was a live alternative. Credence 0.4 is the probability the specific range at equal quality is accurate for the task inventory as a whole. What would change the verdict: validation studies tying exposure ratings to measured speedups (either way), or replicated field trials across diverse occupations showing consistent halving of task times at equal quality.

Decomposition

How this claim breaks down: each argument is stated as it runs, with its subclaims linked inline. ↗︎ opens a subclaim; the map shows how they fit together.

argumentTask-exposure rating estimateThis argument, if it holds, bears in favour of the claim.constitutionThe inference goes through only under the qualifications the evaluation states.constitution

Human annotators and GPT-4 rated the roughly 19,000 tasks in the O*NET database against a rubric asking whether LLM-powered software could cut completion time by at least half at equal quality, and 47 to 56 percent of tasks met that bar once complementary tooling was counted. Because these exposure ratings reliably identify tasks where LLM tools could halve completion time at equal quality, and given that O*NET task descriptions adequately represent the work US workers actually perform and that exposure measures technical potential rather than realized impact, the rated share is a fair estimate of the share of US worker tasks that could be completed significantly faster.

The inference goes through only as far as its rating premise does: the argument stands or falls with whether the exposure ratings reliably identify tasks where LLM tools could halve completion time at equal quality, which is unvalidated against measured speedups and whose absolute calibration shifts markedly with the scoring model. The background premises, that O*NET adequately represents actual work and that exposure measures potential rather than realized impact, are reasonable approximations that qualify precision rather than direction. Granting all three, the argument delivers the range as a rating-based estimate, not as a measured quantity.

argumentExperimental speedup evidenceThis argument, if it holds, bears in favour of the claim.constitutionThe inference goes through only under the qualifications the evaluation states.constitution

Because generative AI assistance substantially reduces task completion time while maintaining or improving quality in controlled experiments on exposed task types such as professional writing and coding, the speedup potential the exposure ratings ascribe to a large share of tasks is demonstrated to be achievable in at least some of the task families the estimate covers.

The argument shows the claimed potential is real somewhere, not that it extends to roughly half of all US worker tasks: demonstrations on self-contained writing and coding tasks establish existence, and generalizing from them is the step the argument cannot supply. It rests on generative AI substantially reducing completion time at equal or better quality, which the experimental record supports in some domains and field evidence undercuts in others.

argumentReal-world slowdown evidenceThis argument, if it holds, weighs against the claim.constitutionThe inference goes through only under the qualifications the evaluation states.constitution

Software development tasks sit near the top of tooling-inclusive exposure ratings, yet using AI tools caused experienced open-source developers to complete real tasks more slowly in early 2025. If rated exposure fails to translate into faster completion in one of the most exposed task families, the rated share overstates the share of tasks that could actually be completed significantly faster at equal quality.

As a counterexample the argument is real but bounded: the finding that early-2025 AI tools slowed experienced open-source developers shows rated exposure failing to convert into speedups in a maximally exposed domain, which weakens confidence in the rated range. It concerns one expert population, familiar codebases, and one tool generation, so it qualifies the parent's range rather than refuting the potential claim outright, and later tools or other populations could show different results.

See how these fit together on the map

or create a grant for this whole area →

Provenance

Where this claim has been said, linked to its canonical form.

When incorporating software and tooling built on top of LLMs, this share increases to between 47 and 56% of all tasks.

Follows the 15% direct-LLM estimate; the paper emphasizes the gap between model and tooling exposure.

Cite this claim: a formal citation with its evidence attached

Contribute

Every judgment on this page is open to challenge. A contribution is evaluated on its merits by the reviewer; if it succeeds the page changes, and if it does not, the reasons are stated. Either way the exchange becomes part of the claim’s public record.


The attention this claim received was paid for by a funded mandate. Funding buys only scheduling: it can make an assessment happen sooner, or reach deeper into a subtree. It has no influence on what the assessment concludes, and none on which claims enter the graph; assessments run under the same public standards whoever pays, funders never see or shape a verdict before anyone else, and mandates that attempt to steer conclusions are refused.

Created by extractor · Aug 11, 2026. Every judgment on this page is accompanied by a reasoning trace.