Browse
Claims
Search the graph by meaning. Each result carries its current verdict; open one to see its decomposition, provenance, and the reasoning behind the assessment.
ShowingImportancePrizesTopicLarge language models
Claims about large language models (LLMs) — neural network models trained on large text corpora to generate and understand human language, including their capabilities, properties, behavior, and impact.
ChatGPT 5.6 Sol Pro produced a counterexample to the Gaussian Moments Conjecture without human intervention after a single initial prompt
UnassessedNo current assessment. Attention goes where its expected value is highest and someone funds it; nothing has funded an assessment of this claim yet, and anyone can.constitution →empirical · derivedA factual claim that rests on inference from other evidence rather than direct observation.constitution →AI mathematical discoveryGaussian Moments Conjectureimportance · settledImportance 0.20, from 0 to 1 · settled: uncontested, so low even when much depends on it. Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution →
Number theory, combinatorics, and graph theory are more accessible to large language models than other areas of mathematics.
ContestedCredible evidence or argument exists on multiple sides.constitution →empirical · derivedA factual claim that rests on inference from other evidence rather than direct observation.constitution →Graph theoryNumber theoryCombinatoricsimportance · minorImportance 0.40, from 0 to 1 · minor: narrow or largely settled, cheap to get right. Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution →
Large language models perform better on elementary, self-contained mathematical problems than on problems requiring extensive theoretical background.
UnassessedNo current assessment. Attention goes where its expected value is highest and someone funds it; nothing has funded an assessment of this claim yet, and anyone can.constitution →empirical · derivedA factual claim that rests on inference from other evidence rather than direct observation.constitution →AI mathematical discoveryMachine insight versus memoryimportance · minorImportance 0.35, from 0 to 1 · minor: narrow or largely settled, cheap to get right. Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution →
Most early AI-driven Erdős problem solutions came from hobbyists using public LLMs rather than corporate labs.
SupportedEvidence favors the claim, but the chain is incomplete or the sources are secondary.constitution →empirical · verifiableA factual claim that could be checked directly against observation or primary records.constitution →AI-assisted Erdős problem solvingAI mathematical discoveryimportance · minorImportance 0.35, from 0 to 1 · minor: narrow or largely settled, cheap to get right. Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution →
Frontier AI models produce novel mathematical results through reasoning comparable to a human mathematician's rather than brute-force search
ContestedCredible evidence or argument exists on multiple sides.constitution →empirical · derivedA factual claim that rests on inference from other evidence rather than direct observation.constitution →AI mathematical discoveryMachine insight versus memoryArtificial intelligenceimportance · notableImportance 0.55, from 0 to 1 · notable: a contested point in a live debate (also the default before judging). Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution →
Reasoning and search are meaningfully distinct modes of problem solving in large language models
UnassessedNo current assessment. Attention goes where its expected value is highest and someone funds it; nothing has funded an assessment of this claim yet, and anyone can.constitution →empirical · derivedA factual claim that rests on inference from other evidence rather than direct observation.constitution →Reasoning versus search in LLMsMachine insight versus memoryimportance · minorImportance 0.40, from 0 to 1 · minor: narrow or largely settled, cheap to get right. Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution →
LLMs find existential conjectures easier to prove than universal "for all" conjectures
UnassessedNo current assessment. Attention goes where its expected value is highest and someone funds it; nothing has funded an assessment of this claim yet, and anyone can.constitution →evaluativeA judgment of worth or quality against some standard: good, fair, effective.constitution →Existential versus universal conjecturesAI mathematical discoveryimportance · minorImportance 0.40, from 0 to 1 · minor: narrow or largely settled, cheap to get right. Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution →
AI language models can find counterexamples to long-standing open mathematical conjectures.
UnassessedNo current assessment. Attention goes where its expected value is highest and someone funds it; nothing has funded an assessment of this claim yet, and anyone can.constitution →evaluativeA judgment of worth or quality against some standard: good, fair, effective.constitution →AI mathematical discoveryArtificial intelligenceimportance · notableImportance 0.60, from 0 to 1 · notable: a contested point in a live debate (also the default before judging). Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution →
With software and tooling built on LLMs, 47 to 56 percent of US worker tasks could be completed significantly faster at equal quality.
ContestedCredible evidence or argument exists on multiple sides.constitution →empirical · derivedA factual claim that rests on inference from other evidence rather than direct observation.constitution →AI task exposure measuresAI Productivity Impactimportance · notableImportance 0.60, from 0 to 1 · notable: a contested point in a live debate (also the default before judging). Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution →
Large language models exhibit the traits of a general-purpose technology
SupportedEvidence favors the claim, but the chain is incomplete or the sources are secondary.constitution →empirical · derivedA factual claim that rests on inference from other evidence rather than direct observation.constitution →General-purpose technologyimportance · notableImportance 0.60, from 0 to 1 · notable: a contested point in a live debate (also the default before judging). Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution →
Headline estimates of LLM exposure at the half-of-tasks threshold include exposure via software and tooling built on LLMs, not LLM access alone
UnassessedNo current assessment. Attention goes where its expected value is highest and someone funds it; nothing has funded an assessment of this claim yet, and anyone can.constitution →empirical · derivedA factual claim that rests on inference from other evidence rather than direct observation.constitution →AI task exposure measuresLLM application layerimportance · minorImportance 0.35, from 0 to 1 · minor: narrow or largely settled, cheap to get right. Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution →
Large language models spawn complementary innovations built on top of them
UnassessedNo current assessment. Attention goes where its expected value is highest and someone funds it; nothing has funded an assessment of this claim yet, and anyone can.constitution →empirical · derivedA factual claim that rests on inference from other evidence rather than direct observation.constitution →LLM application layerGeneral-purpose technologyimportance · minorImportance 0.40, from 0 to 1 · minor: narrow or largely settled, cheap to get right. Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution →
Large language model capabilities are improving rapidly over time
UnassessedNo current assessment. Attention goes where its expected value is highest and someone funds it; nothing has funded an assessment of this claim yet, and anyone can.constitution →empirical · derivedA factual claim that rests on inference from other evidence rather than direct observation.constitution →LLM capability trendsimportance · notableImportance 0.50, from 0 to 1 · notable: a contested point in a live debate (also the default before judging). Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution →
Large language models are applicable to tasks across a wide range of occupations and industries
UnassessedNo current assessment. Attention goes where its expected value is highest and someone funds it; nothing has funded an assessment of this claim yet, and anyone can.constitution →empirical · derivedA factual claim that rests on inference from other evidence rather than direct observation.constitution →AI task exposure measuresimportance · notableImportance 0.50, from 0 to 1 · notable: a contested point in a live debate (also the default before judging). Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution →
AI coding tools became substantially more capable during 2025 with the adoption of agentic tools.
UnassessedNo current assessment. Attention goes where its expected value is highest and someone funds it; nothing has funded an assessment of this claim yet, and anyone can.constitution →empirical · derivedA factual claim that rests on inference from other evidence rather than direct observation.constitution →LLM capability trendsAgentic coding toolsimportance · minorImportance 0.30, from 0 to 1 · minor: narrow or largely settled, cheap to get right. Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution →
Software built on top of LLMs will substantially scale the economic impact of the underlying models
UnassessedNo current assessment. Attention goes where its expected value is highest and someone funds it; nothing has funded an assessment of this claim yet, and anyone can.constitution →causalA claim that one thing brings about another, not merely that the two go together.constitution →LLM application layerimportance · notableImportance 0.50, from 0 to 1 · notable: a contested point in a live debate (also the default before judging). Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution →
AI coding benchmark scores are inflated by data contamination and memorization.
UnassessedNo current assessment. Attention goes where its expected value is highest and someone funds it; nothing has funded an assessment of this claim yet, and anyone can.constitution →empirical · derivedA factual claim that rests on inference from other evidence rather than direct observation.constitution →Benchmark data contaminationAI Coding Benchmarksimportance · notableImportance 0.50, from 0 to 1 · notable: a contested point in a live debate (also the default before judging). Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution →
Contribute
If a claim here is wrong, or missing evidence, open it: every claim page carries its own entry for challenges, evidence, and corrections. If the graph is missing a claim entirely, propose it below. A proposal is reviewed on its merits; accepted claims are matched against the graph and enter it with their reasoning on record.