Minerval

Browse

Claims

Search the graph by meaning. Each result carries its current verdict; open one to see its decomposition, provenance, and the reasoning behind the assessment.

ShowingImportancePrizesTopicLarge language models17 claims

Claims about large language models (LLMs) — neural network models trained on large text corpora to generate and understand human language, including their capabilities, properties, behavior, and impact.

ChatGPT 5.6 Sol Pro produced a counterexample to the Gaussian Moments Conjecture without human intervention after a single initial prompt
UnassessedNo current assessment. Attention goes where its expected value is highest and someone funds it; nothing has funded an assessment of this claim yet, and anyone can.constitutionempirical · derivedA factual claim that rests on inference from other evidence rather than direct observation.constitutionAI mathematical discoveryGaussian Moments Conjectureimportance · settledImportance 0.20, from 0 to 1 · settled: uncontested, so low even when much depends on it. Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution
Number theory, combinatorics, and graph theory are more accessible to large language models than other areas of mathematics.
ContestedCredible evidence or argument exists on multiple sides.constitutionempirical · derivedA factual claim that rests on inference from other evidence rather than direct observation.constitutionGraph theoryNumber theoryCombinatoricsimportance · minorImportance 0.40, from 0 to 1 · minor: narrow or largely settled, cheap to get right. Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution
Large language models perform better on elementary, self-contained mathematical problems than on problems requiring extensive theoretical background.
UnassessedNo current assessment. Attention goes where its expected value is highest and someone funds it; nothing has funded an assessment of this claim yet, and anyone can.constitutionempirical · derivedA factual claim that rests on inference from other evidence rather than direct observation.constitutionAI mathematical discoveryMachine insight versus memoryimportance · minorImportance 0.35, from 0 to 1 · minor: narrow or largely settled, cheap to get right. Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution
Most early AI-driven Erdős problem solutions came from hobbyists using public LLMs rather than corporate labs.
SupportedEvidence favors the claim, but the chain is incomplete or the sources are secondary.constitutionempirical · verifiableA factual claim that could be checked directly against observation or primary records.constitutionAI-assisted Erdős problem solvingAI mathematical discoveryimportance · minorImportance 0.35, from 0 to 1 · minor: narrow or largely settled, cheap to get right. Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution
Frontier AI models produce novel mathematical results through reasoning comparable to a human mathematician's rather than brute-force search
ContestedCredible evidence or argument exists on multiple sides.constitutionempirical · derivedA factual claim that rests on inference from other evidence rather than direct observation.constitutionAI mathematical discoveryMachine insight versus memoryArtificial intelligenceimportance · notableImportance 0.55, from 0 to 1 · notable: a contested point in a live debate (also the default before judging). Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution
Reasoning and search are meaningfully distinct modes of problem solving in large language models
UnassessedNo current assessment. Attention goes where its expected value is highest and someone funds it; nothing has funded an assessment of this claim yet, and anyone can.constitutionempirical · derivedA factual claim that rests on inference from other evidence rather than direct observation.constitutionReasoning versus search in LLMsMachine insight versus memoryimportance · minorImportance 0.40, from 0 to 1 · minor: narrow or largely settled, cheap to get right. Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution
LLMs find existential conjectures easier to prove than universal "for all" conjectures
UnassessedNo current assessment. Attention goes where its expected value is highest and someone funds it; nothing has funded an assessment of this claim yet, and anyone can.constitutionevaluativeA judgment of worth or quality against some standard: good, fair, effective.constitutionExistential versus universal conjecturesAI mathematical discoveryimportance · minorImportance 0.40, from 0 to 1 · minor: narrow or largely settled, cheap to get right. Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution
AI language models can find counterexamples to long-standing open mathematical conjectures.
UnassessedNo current assessment. Attention goes where its expected value is highest and someone funds it; nothing has funded an assessment of this claim yet, and anyone can.constitutionevaluativeA judgment of worth or quality against some standard: good, fair, effective.constitutionAI mathematical discoveryArtificial intelligenceimportance · notableImportance 0.60, from 0 to 1 · notable: a contested point in a live debate (also the default before judging). Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution
With software and tooling built on LLMs, 47 to 56 percent of US worker tasks could be completed significantly faster at equal quality.
ContestedCredible evidence or argument exists on multiple sides.constitutionempirical · derivedA factual claim that rests on inference from other evidence rather than direct observation.constitutionAI task exposure measuresAI Productivity Impactimportance · notableImportance 0.60, from 0 to 1 · notable: a contested point in a live debate (also the default before judging). Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution
Large language models exhibit the traits of a general-purpose technology
SupportedEvidence favors the claim, but the chain is incomplete or the sources are secondary.constitutionempirical · derivedA factual claim that rests on inference from other evidence rather than direct observation.constitutionGeneral-purpose technologyimportance · notableImportance 0.60, from 0 to 1 · notable: a contested point in a live debate (also the default before judging). Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution
Headline estimates of LLM exposure at the half-of-tasks threshold include exposure via software and tooling built on LLMs, not LLM access alone
UnassessedNo current assessment. Attention goes where its expected value is highest and someone funds it; nothing has funded an assessment of this claim yet, and anyone can.constitutionempirical · derivedA factual claim that rests on inference from other evidence rather than direct observation.constitutionAI task exposure measuresLLM application layerimportance · minorImportance 0.35, from 0 to 1 · minor: narrow or largely settled, cheap to get right. Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution
Large language models spawn complementary innovations built on top of them
UnassessedNo current assessment. Attention goes where its expected value is highest and someone funds it; nothing has funded an assessment of this claim yet, and anyone can.constitutionempirical · derivedA factual claim that rests on inference from other evidence rather than direct observation.constitutionLLM application layerGeneral-purpose technologyimportance · minorImportance 0.40, from 0 to 1 · minor: narrow or largely settled, cheap to get right. Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution
Large language model capabilities are improving rapidly over time
UnassessedNo current assessment. Attention goes where its expected value is highest and someone funds it; nothing has funded an assessment of this claim yet, and anyone can.constitutionempirical · derivedA factual claim that rests on inference from other evidence rather than direct observation.constitutionLLM capability trendsimportance · notableImportance 0.50, from 0 to 1 · notable: a contested point in a live debate (also the default before judging). Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution
Large language models are applicable to tasks across a wide range of occupations and industries
UnassessedNo current assessment. Attention goes where its expected value is highest and someone funds it; nothing has funded an assessment of this claim yet, and anyone can.constitutionempirical · derivedA factual claim that rests on inference from other evidence rather than direct observation.constitutionAI task exposure measuresimportance · notableImportance 0.50, from 0 to 1 · notable: a contested point in a live debate (also the default before judging). Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution
AI coding tools became substantially more capable during 2025 with the adoption of agentic tools.
UnassessedNo current assessment. Attention goes where its expected value is highest and someone funds it; nothing has funded an assessment of this claim yet, and anyone can.constitutionempirical · derivedA factual claim that rests on inference from other evidence rather than direct observation.constitutionLLM capability trendsAgentic coding toolsimportance · minorImportance 0.30, from 0 to 1 · minor: narrow or largely settled, cheap to get right. Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution
Software built on top of LLMs will substantially scale the economic impact of the underlying models
UnassessedNo current assessment. Attention goes where its expected value is highest and someone funds it; nothing has funded an assessment of this claim yet, and anyone can.constitutioncausalA claim that one thing brings about another, not merely that the two go together.constitutionLLM application layerimportance · notableImportance 0.50, from 0 to 1 · notable: a contested point in a live debate (also the default before judging). Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution
AI coding benchmark scores are inflated by data contamination and memorization.
UnassessedNo current assessment. Attention goes where its expected value is highest and someone funds it; nothing has funded an assessment of this claim yet, and anyone can.constitutionempirical · derivedA factual claim that rests on inference from other evidence rather than direct observation.constitutionBenchmark data contaminationAI Coding Benchmarksimportance · notableImportance 0.50, from 0 to 1 · notable: a contested point in a live debate (also the default before judging). Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution

Contribute

If a claim here is wrong, or missing evidence, open it: every claim page carries its own entry for challenges, evidence, and corrections. If the graph is missing a claim entirely, propose it below. A proposal is reviewed on its merits; accepted claims are matched against the graph and enter it with their reasoning on record.