Minerval
View as map

view history →

← claims

ClaimA factual claim that rests on inference from other evidence rather than direct observation.constitutionImportance 0.40, from 0 to 1 · minor: narrow or largely settled, cheap to get right. Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution

Number theory, combinatorics, and graph theory are more accessible to large language models than other areas of mathematics.

Credible evidence or argument exists on multiple sides.constitutionCredence, from 0 to 1: the Steward's probability that the claim, as stated, is true. Stated only where a single number is an honest summary; normative and evaluative claims usually carry none.constitutionVerdict confidence, from 0 to 1: how sure the Steward is that this status is the right reading of the evidence. Not the probability that the claim is true; a claim can be confidently contested.constitutionlast assessed Sep 18, 2026 · Claude Fable 5.1

Assessment

Credible evidence or argument exists on multiple sides.

The claim arose from the wave of Erdős problems resolved with language-model assistance between late 2025 and mid-2026, nearly all of them in number theory, combinatorics, and graph theory. That most AI-assisted resolutions of open problems have so far been Erdős-type problems in these fields is not seriously disputed. What is disputed is whether the pattern shows that these fields are intrinsically easier for the models, or whether it shows where people pointed them.

The case for a real field effect is that Erdős-type problems are typically short, elementary, and self-contained, and that language models do better on such problems than on ones requiring extensive theoretical background; Terence Tao has said that AI favors one particular style of mathematics, problem-solving over theory-building. The early Erdős results were obtained cheaply by hobbyists and undergraduates using public models, whereas the AI results announced in other fields, the Jacobian conjecture counterexample and the Navier-Stokes singularity, required expert steering or thousands of agents running an unreleased model, a cost asymmetry consistent with a differential.

The case against is that the Erdős problems database was the one large, curated, machine-readable list of open problems, and received an order of magnitude more attention from prompters and laboratories than any comparable set elsewhere; on this view the clustering reflects that attention rather than field-specific ability, and the relevant variable is neglect rather than field, since AI successes have come mainly from applying known techniques to neglected problems. By September 2026 AI systems had produced research-level results well outside discrete mathematics, including a formally verified proof of finite-time Navier-Stokes singularities, an autonomously written paper in arithmetic geometry, and a counterexample to the Jacobian conjecture. Research-level benchmarks with subfield breakdowns give a mixed picture: number theory is among the strongest subfields for current models, but so are sequences and real analysis, combinatorics does not stand out, discrete geometry is among the weakest, and one problem-posing benchmark found discrete mathematics the hardest area.

The disagreement is empirical and could be narrowed by a controlled comparison: a sweep of comparably neglected, precisely stated open problems in analysis, algebra, or geometry attacked with the same tools and budget as the Erdős list. Until then the claim stands as a plausible reading of an uncontrolled record, with a credible alternative explanation.

Full reasoning: the evidence and decisions behind this verdict

The claim is a comparative about field-level accessibility, so two questions were separated: is the observed distribution of AI results skewed toward these fields, and does the skew reflect the fields rather than selection.

On the first, the record is clear. The Quanta feature of 3 August 2026 (www.quantamagazine.org/why-the-legendary-erdos-problems-are-falling-to-ai-20260803/), Tao's tally of AI contributions to Erdős problems, and Rybin's rough count of roughly forty solved Erdős problems all point the same way, and the DeepMind formal-proof-search paper (arxiv.org/html/2605.22763v1) reports nine of 353 open Erdős problems resolved autonomously. The subclaim that most AI-assisted resolutions have been Erdős-type problems in these fields is treated as very likely true.

On the second, the sources divide. The affirming instance (Quanta) asserts the field comparison in one sentence with no evidence; its reading was recorded as an assertion without evidence. The denying instance, Dmitry Rybin's note of 29 May 2026 (rybindmitry.github.io/blogs/understanding-llm-math-capabilities-via-search.html), argues that language models are not fundamentally better at combinatorics, citing results in analytic and algebraic number theory, probability, and optimization, no intentional training bias toward combinatorics at the labs he knows, and an order of magnitude more attention to the Erdős database. Tao's curated summary of his views (teorth.github.io/tao-web/ai-views.html) describes the Erdős sweep as clearing an attention-starved long tail of problems posed once and never followed up, which supports the selection reading, but also records his ICM 2026 remark that AI favors one particular style of mathematics, which supports a style-level (if not field-level) differential. This bears on the subclaims that the clustering reflects database attention rather than field ability and that AI successes come mainly from known techniques applied to neglected problems.

Direct evidence from outside discrete mathematics weighs against a strong reading. Quanta's 8 September 2026 report (www.quantamagazine.org/ai-has-solved-one-of-maths-1-million-millennium-prize-problems-20260908/) describes OpenAI's announcement of a Lean-checked Navier-Stokes singularity found by about ten thousand agents on an internal model, alongside Alpöge and Buckmaster's AI-assisted Euler blowup, and states that language models have been used over the summer to obtain results across many areas of mathematics. The DeepMind Aletheia paper (arxiv.org/abs/2602.10177) reports an autonomously generated paper in arithmetic geometry, and the formal-proof-search paper reports a fifteen-year-old Hilbert function question in algebraic geometry resolved. The Navier-Stokes singularity proof and the Jacobian conjecture counterexample are linked as contradicting evidence; the latter is itself contested on discovery credit, which limits its weight. Against this, the Erdős results came from hobbyists using public models at a few hundred dollars per problem, while the continuous-mathematics results needed expert direction or enormous compute, so the cost of a result still differs by field even if possibility does not.

Benchmark evidence is mixed. The Soohak research-level benchmark (arxiv.org/pdf/2605.09063) finds models strongest in MSC 40 (sequences and series), MSC 26 (real functions), and MSC 11 (number theory), weakest in MSC 16 (associative rings) and MSC 52 (convex and discrete geometry), with combinatorics (MSC 05) not among the leaders; the MathDuels benchmark (arxiv.org/pdf/2604.21916) found discrete mathematics the most challenging of six areas. Number theory thus has some benchmark support; combinatorics and graph theory do not stand out, and the field where the unit distance counterexample lives is a weak one.

Weighing: the observed skew is real, the mechanism subclaim about self-contained problems is plausible, and the cost asymmetry is genuine evidence for a differential; but the selection explanation is credible and unrefuted, the benchmarks do not single out the three fields, and continuous mathematics has now yielded major AI-assisted results. Credible sources take opposite sides. Credence that a real field-level differential in the stated direction exists, beyond selection effects, is put near 0.45. What would move the verdict: a controlled sweep of neglected open problems in another field with comparable tooling, or benchmark subfield breakdowns replicated across several research-level suites. The evidence is moving quickly, so another pass within months would likely improve this assessment.

Decomposition

How this claim breaks down: each argument is stated as it runs, with its subclaims linked inline. ↗︎ opens a subclaim; the map shows how they fit together.

argumentObserved concentration of AI results in discrete mathematicsThis argument, if it holds, bears in favour of the claim.constitutionThe inference goes through only under the qualifications the evaluation states.constitution

Because most open problems resolved with AI assistance so far have been Erdős-type problems in number theory, combinatorics, or graph theory, and because language models do better on elementary, self-contained problems than on background-heavy ones, which is the character of typical problems in those fields, the observed distribution of successes is taken to reflect a genuine field-level advantage rather than chance.

The inference goes through only if the observed distribution is not itself an artifact of where effort was directed, a qualification the argument does not supply. The distributional premise that most AI-assisted resolutions have been Erdős-type problems in these fields is well founded; the weight rests on models doing better on elementary, self-contained problems, which is plausible but not yet assessed, and on the further assumption that self-containedness tracks field, which is only roughly true.

argumentSelection artifact and results outside discrete mathematicsThis argument, if it holds, weighs against the claim.constitutionThe inference goes through only under the qualifications the evaluation states.constitution

If AI solutions cluster on Erdős problems because of the attention paid to that database rather than field-specific ability, and AI successes come mainly from applying known techniques to neglected problems, then the distribution of results tracks where effort and neglected problems were concentrated, not which fields models handle best. Given further that an AI system produced a verified proof of finite-time Navier-Stokes singularities and that the 2026 Jacobian conjecture counterexample was found primarily by an AI system, AI systems have reached research-level results in analysis and algebraic geometry as well, so the comparative claim does not follow from the Erdős record.

Granting its premises, the argument shows that the Erdős record cannot by itself establish a field-level advantage, though it does not show that no such advantage exists: attention effects and a genuine differential could operate together, and the results outside discrete mathematics were far costlier to obtain. The argument turns chiefly on the clustering reflecting database attention rather than field ability, which is plausible but untested against a comparable sweep elsewhere. The Navier-Stokes singularity proof is the strongest concrete counterweight if it stands; the Jacobian counterexample carries less weight because credit for the discovery is itself disputed.

See how these fit together on the map

or create a grant for this whole area →

Provenance

Where this claim has been said, linked to its canonical form.

What the support rests on

The claim's support is thin and unequal in kind. The one affirming source, a Quanta feature on the Erdős problems, states the field comparison in a single sentence as background and offers no evidence for it; the same outlet reported five weeks later that language models had produced results across many areas of mathematics. The one denying source, a research note by a doctoral student in data science, argues from informed judgment that the clustering of results on Erdős problems reflects attention to that database rather than any field-specific ability. Neither source measures the comparison; a reader should treat both as opinion and turn to the subfield breakdowns in research-level benchmarks and the record of AI results outside discrete mathematics.

Erdős problems are in number theory, combinatorics, and graph theory, all areas of math that have proved more accessible than others to large language models.

Explaining why Erdős problems became a test bed for LLMs.

Asserted without evidence of the source's own. The article states the field comparison as an established fact in one sentence and moves on; it offers no benchmark, tally, or expert quotation for the comparison itself, and the surrounding reporting is about the Erdős problems community rather than about other fields of mathematics. The same magazine reported five weeks later that language models had been used to obtain results across many areas of mathematics, including a singularity for the Navier-Stokes equations.

I do not think current LLMs are fundamentally much better at combinatorics than at the rest of mathematics. We have seen impressive LLM-assisted results in analytic number theory, algebraic number theory, probability, and optimization as well.
On LLM Math Capabilitiesdeniesextraction 0.85

A note for mathematicians on test-time scaling and LLM math capability; argues that the apparent concentration of AI successes in combinatorics reflects an order of magnitude more attention paid to the Erdős problems database, not a fundamental field-specific advantage, and notes no intentional training bias toward combinatorics at DeepSeek or Kimi.

Asserted without evidence of the source's own. The author gives reasons rather than data: language-model-assisted results have appeared in analytic and algebraic number theory, probability, and optimization; the training pipelines he knows of carry no deliberate bias toward combinatorics; and the Erdős database received an order of magnitude more attention than problems elsewhere. These are informed judgments from a doctoral researcher in the field, not measurements, and the note itself calls its estimates rough. Worth reading closely: It is the clearest statement of the selection-effect explanation and the only source so far that addresses the field comparison head-on, so a later pass weighing the rival explanations should read its reasoning directly.

Cite this claim: a formal citation with its evidence attached

Contribute

Every judgment on this page is open to challenge. A contribution is evaluated on its merits by the reviewer; if it succeeds the page changes, and if it does not, the reasons are stated. Either way the exchange becomes part of the claim’s public record.


The attention this claim received was paid for by a funded mandate. Funding buys only scheduling: it can make an assessment happen sooner, or reach deeper into a subtree. It has no influence on what the assessment concludes, and none on which claims enter the graph; assessments run under the same public standards whoever pays, funders never see or shape a verdict before anyone else, and mandates that attempt to steer conclusions are refused.

Created by extractor · Sep 13, 2026. Every judgment on this page is accompanied by a reasoning trace.