Minerval

Browse

Claims

Search the graph by meaning. Each result carries its current verdict; open one to see its decomposition, provenance, and the reasoning behind the assessment.

ShowingImportancePrizesTopicLLM capability trends7 claims

Claims about the trajectory or pace of improvement in large language model capabilities over time — e.g., benchmark gains across model generations, scaling-driven gains, and whether progress is accelerating, steady, or plateauing. Excludes compute/cost trends (covered by Training compute trends) and forecasts of human-level AI timelines.

The 2026 AI mathematical results mark a phase transition in AI models' research mathematics capability.
SupportedEvidence favors the claim, but the chain is incomplete or the sources are secondary.constitutionevaluativeA judgment of worth or quality against some standard: good, fair, effective.constitutionAI mathematical discoveryArtificial intelligenceimportance · majorImportance 0.65, from 0 to 1 · major: real consequence within a domain, actively argued. Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution
Rapid AI capability improvements will extend productivity gains to hard-to-learn tasks within the next decade.
ContestedCredible evidence or argument exists on multiple sides.constitutionempirical · derivedA factual claim that rests on inference from other evidence rather than direct observation.constitutionAI Productivity ImpactAI and skill gapsimportance · notableImportance 0.60, from 0 to 1 · notable: a contested point in a live debate (also the default before judging). Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution
Scaling up training compute will remain the primary driver of frontier AI capability gains
UnassessedNo current assessment. Attention goes where its expected value is highest and someone funds it; nothing has funded an assessment of this claim yet, and anyone can.constitutionempirical · derivedA factual claim that rests on inference from other evidence rather than direct observation.constitutionArtificial intelligenceimportance · notableImportance 0.60, from 0 to 1 · notable: a contested point in a live debate (also the default before judging). Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution
AI occupational exposure scores are snapshots of a specific model's capabilities that become outdated as AI capabilities advance
UnassessedNo current assessment. Attention goes where its expected value is highest and someone funds it; nothing has funded an assessment of this claim yet, and anyone can.constitutionempirical · derivedA factual claim that rests on inference from other evidence rather than direct observation.constitutionAI task exposure measuresimportance · minorImportance 0.30, from 0 to 1 · minor: narrow or largely settled, cheap to get right. Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution
The complexity of tasks AI systems can complete autonomously has been growing rapidly year over year.
UnassessedNo current assessment. Attention goes where its expected value is highest and someone funds it; nothing has funded an assessment of this claim yet, and anyone can.constitutionempirical · derivedA factual claim that rests on inference from other evidence rather than direct observation.constitutionArtificial intelligenceimportance · notableImportance 0.50, from 0 to 1 · notable: a contested point in a live debate (also the default before judging). Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution
Large language model capabilities are improving rapidly over time
UnassessedNo current assessment. Attention goes where its expected value is highest and someone funds it; nothing has funded an assessment of this claim yet, and anyone can.constitutionempirical · derivedA factual claim that rests on inference from other evidence rather than direct observation.constitutionLarge language modelsimportance · notableImportance 0.50, from 0 to 1 · notable: a contested point in a live debate (also the default before judging). Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution
AI coding tools became substantially more capable during 2025 with the adoption of agentic tools.
UnassessedNo current assessment. Attention goes where its expected value is highest and someone funds it; nothing has funded an assessment of this claim yet, and anyone can.constitutionempirical · derivedA factual claim that rests on inference from other evidence rather than direct observation.constitutionLarge language modelsAgentic coding toolsimportance · minorImportance 0.30, from 0 to 1 · minor: narrow or largely settled, cheap to get right. Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution

Contribute

If a claim here is wrong, or missing evidence, open it: every claim page carries its own entry for challenges, evidence, and corrections. If the graph is missing a claim entirely, propose it below. A proposal is reviewed on its merits; accepted claims are matched against the graph and enter it with their reasoning on record.