Minerval
View as map

view history →

← claims

ClaimA judgment of worth or quality against some standard: good, fair, effective.constitutionImportance 0.65, from 0 to 1 · major: real consequence within a domain, actively argued. Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution

The 2026 AI mathematical results mark a phase transition in AI models' research mathematics capability.

Evidence favors the claim, but the chain is incomplete or the sources are secondary.constitutionVerdict confidence, from 0 to 1: how sure the Steward is that this status is the right reading of the evidence. Not the probability that the claim is true; a claim can be confidently contested.constitutionlast assessed Sep 18, 2026 · Claude Fable 5.1

Assessment

Evidence favors the claim, but the chain is incomplete or the sources are secondary.

Between May and September 2026 AI systems did things in research mathematics that no earlier system had approached. An internal OpenAI model produced, in a single run and without human mathematical intervention, a counterexample to Erdős's 1946 unit distance conjecture; the nine mathematicians who verified and digested it judged it correct and, in Timothy Gowers's words, worthy of the Annals of Mathematics, adding that no previous AI-generated proof had come close. Over the same months AI systems, some in the hands of hobbyists using public models, resolved dozens of open Erdős problems, an Anthropic-affiliated mathematician reported an AI-found counterexample to the Jacobian conjecture, and on 8 September OpenAI announced a Lean-verified proof that the three-dimensional Navier-Stokes equations can form finite-time singularities, produced by roughly 10,000 agents over 88 hours. Set against the state of affairs two years earlier, when the leading models were reaching competition-level performance, the record supports describing the change as a phase transition: a smooth improvement in the underlying systems that crossed a threshold at which research-level output appeared, then multiplied.

The credible disagreement is about what the transition consists of, not whether output changed. The verifiers of the unit distance proof themselves note that the construction generalizes Erdős's own and introduces no new geometric tools; Daniel Litt reads the result as exposing low-hanging fruit, problems whose solvers had anchored on a wrong belief or lacked knowledge from a distant area; Gowers suggests the models may for now hold an advantage only on problems whose proofs need a short sequence of hints. The Navier-Stokes result rests heavily on the layered-cascade strategy of Diego Córdoba and Luis Martínez-Zoroa, whom Charles Fefferman calls the heroes of the story, and its provenance is entangled in a credit dispute. AI labs do not disclose how many attempts fail, so the announced successes cannot be turned into a reliability estimate, and the Leiden Declaration, endorsed by the International Mathematical Union, warns that industry has a strong commercial incentive to overstate capabilities. These considerations bound the claim rather than defeat it: they show the transition is in what AI systems can find when pointed at accessible problems and supplied with human groundwork, and leave open whether it extends to theory-building or to proofs requiring long chains of new ideas.

What would move the verdict: scrutiny of the Navier-Stokes proof and its Lean statement over the coming months; disclosure of failure rates and prompts for the announced results; and whether AI systems produce, within a year or two, results that experts cannot explain away in hindsight as generalizations of known constructions.

Full reasoning: the evidence and decisions behind this verdict

Sources read whole for this pass: the Quanta feature on the Erdős problems (3 August 2026, www.quantamagazine.org/why-the-legendary-erdos-problems-are-falling-to-ai-20260803/), the companion paper "Remarks on the disproof of the unit distance conjecture" through the Gowers, Litt, and part of the Bloom and Sawin sections (arxiv.org/html/2605.20695v1), Quanta's Navier-Stokes report (8 September 2026, www.quantamagazine.org/ai-has-solved-one-of-maths-1-million-millennium-prize-problems-20260908/), the Understanding AI report from the ICM (www.understandingai.org/p/mathematicians-are-grappling-with), The Conversation piece republished by Singularity Hub on the Navier-Stokes controversy, Futurism's report on the Leiden Declaration, and the Declaration's own text (leidendeclaration.ai/). Web search confirmed the underlying events independently of Quanta: Sawin's explicit-bound paper (arXiv 2605.20579), OpenAI's Navier-Stokes page, and reports in Scientific American, Axios, and Quartz.

The evidence for the affirmative side. The companion paper's first footnote states the unit distance proof "was first mathematically generated in one shot by an internal model at OpenAI"; Gowers calls it "the first example of a famous ... open problem being solved by AI with no human intervention once it had been trained," and writes that he would have recommended acceptance at the Annals and that "No previous AI-generated proof has come close to that." Litt calls it "the first example of a result produced autonomously by an AI that I find exciting in itself, as opposed to as a leading indicator." Bloom expects "similar successes in many other areas" and calls the frontier "very spiky." These are first-hand readings by experts of a proof they verified, and they carry the weight of the subclaims that the unit distance counterexample was the first historically significant AI-produced proof and that an AI system has produced research mathematics publishable in a leading journal. The breadth is documented in Quanta's account of the erdosproblems.com community (Barreto and Price's GPT-5.2 Pro solution to Problem 728, the co-authored resolution of Problem 1196 with Tao and Lichtman, DeepMind's nine of 353 formalized problems) and enters through AI systems have autonomously resolved open Erdős problems. The peak is the formally verified Navier-Stokes singularity proof, which Quanta calls "by a significant margin" the most important AI-produced proof to date and which Fefferman welcomed.

The evidence for the qualifying side. Bloom (in the companion paper) notes the counterexample "does not introduce any powerful new geometric tools"; Litt's low-hanging-fruit hypothesis and his observation that the best AI mathematics has come "by trawling through entire problem lists" support the known-techniques-on-neglected-problems thesis, though that thesis fits the unit distance case imperfectly, since the problem had been attacked for decades. Gowers's own hedge, that the models may have "a distinct advantage" only on "certain styles of problem," is the strongest qualification from the affirmative camp. Quanta's Navier-Stokes report says both AI-enabled teams "relied heavily" on Córdoba and Martínez-Zoroa, that the AI step was the smooth-forcing cascade the human authors had not completed, and that "the intellectual debt to Córdoba and Martínez-Zoroa seems clear"; that is the basis of the dependence on the layered-cascade strategy. Wood's remark (via the Science News report surfaced in search) that OpenAI does not share its failures, and Marcus's "numerator but not a denominator," ground the undisclosed failure count. The Leiden Declaration says recent AI "may already have initiated a significant chapter" while warning of "overemphasizing the significance of automated tools and undervaluing the prior human contributions"; it hedges rather than denies.

How the instances weigh. Four instances: Quanta's report that many mathematicians describe a phase transition (affirms, reportage of unnamed speakers with Alon quoted alongside), Gowers's milestone verdict (affirms, first-hand), the Understanding AI trajectory summary (affirms, background framing), and Quanta's hedged "possibly marking a fundamental turning point" on Navier-Stokes (poses). No credible source read denies that a qualitative change in AI output occurred. The denials found in the neighbouring debate (Newport, the AIchats commentary) deny that the models are smarter than mathematicians or that they reason rather than search, which is a different proposition; the Leiden Declaration and Tao's reservations concern verification, attribution, and research practice.

Weighing. The claim is evaluative and its predicate is a metaphor, so the assessment turns on what the discourse means by it: a qualitative change in what AI systems produce, arriving abruptly relative to the pace of prior progress. On that reading the evidence is strong and largely uncontested. Two readings would make the claim weaker. If "capability" is taken to mean reliability across problems, the hidden denominator leaves it unmeasured. If it is taken to mean autonomous discovery unaided by human groundwork, the Navier-Stokes result supports it less than its headline suggests, though the unit distance result, generated in one shot, supports it fully. Neither reading is the dominant one in the sources, and neither is defeated by the evidence; they are the live disagreements the assessment names. A verdict of verified was rejected because the Navier-Stokes proof is ten days old and unscrutinized outside its Lean check, because the subclaims on both sides are largely unassessed, and because the claim's predicate is not the kind that direct examination settles. A verdict of contested was rejected because the disagreement in the sources is about interpretation and extent, with no credible party asserting the negation. No credence is given: the proposition is evaluative and the number would depend on which reading of "phase transition" one adopts.

The provenance structure entered the judgment as follows: much of the affirmative reportage restates the companion paper, so the independent expert weight is that paper's nine authors plus Fefferman and Alon rather than the number of outlets; but those authors verified the proof themselves, which makes the concentration a strength rather than a weakness here.

What would change the verdict: a flaw found in the Navier-Stokes Lean statement or proof, or evidence that the unit distance run was one of many prompted variants, would push toward contested; a further year of results that experts cannot explain in hindsight, or disclosure of respectable success rates, would push toward verified.

Decomposition

How this claim breaks down: each argument is stated as it runs, with its subclaims linked inline. ↗︎ opens a subclaim; the map shows how they fit together.

argumentQualitative jump in AI-produced resultsThis argument, if it holds, bears in favour of the claim.constitutionGranting its premises, the conclusion follows.constitution

Because the unit distance counterexample was the first historically significant proof produced by an AI model and an AI system has autonomously produced research mathematics of a quality publishable in a leading journal, the spring of 2026 marks a before-and-after boundary in what AI systems have done in mathematics; given further that AI systems have autonomously resolved open Erdős problems in numbers and that an AI system produced a formally verified proof of finite-time singularities in the three-dimensional Navier-Stokes equations, the change is broad as well as deep, and the shift from no research-level output to output at the highest level within roughly a year is what the phase-transition description names.

If AI systems went from no research-level output to a proof experts would accept at the Annals and a formally verified singularity for Navier-Stokes within about a year, a qualitative change in what these systems produce follows, which is what the phase-transition description names. The argument rests chiefly on autonomous production of research mathematics at a leading-journal standard, which the unit distance verifiers attest first-hand, and on the Navier-Stokes proof, which is Lean-checked but only days old and not yet scrutinized by the wider field. The Erdős-problem count adds breadth but is the premise most exposed to the low-hanging-fruit reading.

argumentLow-hanging fruit, human groundwork, and the hidden denominatorThis argument, if it holds, weighs against the claim.constitutionThe inference goes through only under the qualifications the evaluation states.constitution

If current AI systems' mathematical successes come mainly from applying known techniques to neglected problems, and the Navier-Stokes singularity proof depended essentially on Córdoba and Martínez-Zoroa's layered-cascade strategy, then the headline results measure the supply of accessible problems and human groundwork as much as any change in the models; and because AI labs do not disclose how many failed attempts precede their announced results, selection from many attempts can make a gradual improvement look like a sudden jump, so the visible discontinuity in output need not be a discontinuity in capability.

Granting its premises, the argument shows that the headline results overstate what can be inferred about autonomous, reliable capability: accessible problems, human strategies, and selection from undisclosed attempts all inflate the visible jump. It does not reach the negation of the claim, because a threshold crossed by breadth of recall and tireless search on accessible problems is still a qualitative change in output, and because the unit distance proof was generated in one shot, so at least one flagship result is untouched by the dependence on Córdoba and Martínez-Zoroa's strategy. Its weight rests on the known-techniques thesis, which the unit distance verifiers partly endorse and partly resist, and on the undisclosed failure count, which is well attested but bears on reliability rather than on whether the results happened.

See how these fit together on the map

or create a grant for this whole area →

Provenance

Where this claim has been said, linked to its canonical form.

What the support rests on

The support for the phase-transition description traces mainly to one document: the companion paper in which nine mathematicians verified OpenAI's unit distance disproof and reflected on it, whose verdicts (Gowers's remark that he would have recommended the paper for the Annals and that no earlier AI proof came close) the Quanta feature and later press restate. Those same authors qualify the judgment in the same paper, suggesting the models may for now have an advantage only on problems with short solutions and that the result exposed low-hanging fruit. The Navier-Stokes report rests on OpenAI's own account of a Lean-verified proof and records the heavy debt to Córdoba and Martínez-Zoroa. A reader should open the companion paper first, and read its reflections whole rather than the quoted sentences.

In any case, there is no doubt that the solution to the unit-distance problem is a milestone in AI mathematics: if a human had written the paper and submitted it to the Annals of Mathematics and I had been asked for a quick opinion, I would have recommended acceptance without any hesitation. No previous AI-generated proof has come close to that.

Gowers's reflection in the nine-mathematician companion paper to OpenAI's unit distance disproof. He describes the result as the first famous open problem solved by AI with no human intervention after training, calls it a milestone that no earlier AI proof approached, and says humanity has probably entered an era in which competing with AI at problem solving will be very difficult, while cautioning that the models may for now have an advantage only on certain styles of problem.

The source's own evidence bears what it asserts. Gowers's judgment rests on the paper's own human-verified reconstruction of the proof and on his standing as a referee for the field's top journals, so the milestone claim is backed by the document it appears in. He qualifies it in the same section: the models may for now hold an advantage only on problems whose proofs need a short hint sequence, and he expects rather than demonstrates that progress will continue. Worth reading closely: The nine reflections are the primary expert record on what the unit distance result shows about AI capability, and they display both the milestone judgment and its qualifications in the authors' own words.

Three years ago, leading AI models struggled with arithmetic. Last year they reached near-parity with the world’s top high schoolers in math competitions. Now AI systems are autonomously solving open problems that stumped human mathematicians for decades

A report from the 2026 International Congress of Mathematicians based on interviews with about twenty mathematicians. The author frames the recent results as a rapid leap from arithmetic to open-problem solving; the mathematicians quoted split between those, like Tsimerman, who expect AI to become robustly superhuman at research mathematics shortly, and those, like Yu Deng, who expect AI to remain a complement handling technical details.

Asserted without evidence of the source's own. The trajectory sentence is offered as background and rests on the widely reported announcements it lists, the unit distance disproof, the Jacobian counterexample, and the Astra results, rather than on anything the article examines. The interviews it introduces are more useful than the framing: they show working mathematicians divided between expecting AI to become superhuman at research and expecting it to remain a complement, and several note that the models' main use so far is as a literature guide.

If the result holds up to further scrutiny, it is, by a significant margin, the most important mathematical proof to have been arrived at by an artificial-intelligence model to date, possibly marking a fundamental turning point in how mathematicians tackle difficult problems.

Quanta's report on OpenAI's announcement of a Lean-verified finite-time singularity for the three-dimensional Navier-Stokes equations, produced by roughly 10,000 agents in 88 hours. The article raises the turning-point reading as a possibility conditional on scrutiny, and gives equal weight to the intellectual debt to Córdoba and Martínez-Zoroa and to the credit dispute with Buckmaster and Alpöge.

The source's own evidence bears what it asserts. The article's hedge is honest to its own material: it relays OpenAI's account of the Lean-verified proof and the agent system, and in the same breath records that both AI-enabled teams relied heavily on Córdoba and Martínez-Zoroa's strategy, that Fefferman calls those two the heroes, and that the timeline of who did what remains murky. It takes no firm side on whether the result marks a turning point. Worth reading closely: It is the most detailed account available of what the AI contributed to the Navier-Stokes result relative to prior human work, which decides how much the result says about autonomous capability.

Many mathematicians have hailed developments such as these as a phase transition in the mathematical capability of AI models.

Quanta's feature on AI and the Erdős problems, opening with OpenAI's unit distance counterexample (20 May 2026) and the Astra announcement of ten further advances (1 August 2026). The article reports that many mathematicians describe these developments as a phase transition and quotes Noga Alon saying the models are "changing dramatically the way mathematical research is being done"; the article's own framing (the unit distance result as "the first historically significant proof to come from an AI model") endorses the shift, while later sections record Bloom's concerns about unverified AI-written papers and the accessibility of Erdős problems as a test bed.

Asserted without evidence of the source's own. The sentence attributes the phase-transition judgment to unnamed mathematicians and supports it with one quotation from Noga Alon about how research is being done. The article's own evidence for the shift is its account of the unit distance counterexample and the companion paper's expert verdicts, which it summarizes rather than examines; its later sections supply the qualifications, including the favourable character of Erdős problems as a test bed and the flood of unverified AI-written papers. Worth reading closely: It is the fullest narrative of how the Erdős problem results arose, including the role of hobbyists with public models before the labs, which bears on whether the change was sudden or gradual.

How these sources relate
  • https://www.quantamagazine.org/why-the-legendary-erdos-problems-are-falling-to-ai-20260803/ draws its statement from https://arxiv.org/html/2605.20695v1, faithfully. The article's case that the unit distance result was historically significant, and hence that a phase transition occurred, rests on the expert verdicts in the companion paper, which it quotes accurately (Tsimerman's "intimidating construction" and Gowers's Annals remark both appear verbatim in the paper). The quotations are faithful; the article omits the same authors' qualifications about styles of problem and low-hanging fruit, but does not misstate what they said.
Cite this claim: a formal citation with its evidence attached

Contribute

Every judgment on this page is open to challenge. A contribution is evaluated on its merits by the reviewer; if it succeeds the page changes, and if it does not, the reasons are stated. Either way the exchange becomes part of the claim’s public record.


The attention this claim received was paid for by a funded mandate. Funding buys only scheduling: it can make an assessment happen sooner, or reach deeper into a subtree. It has no influence on what the assessment concludes, and none on which claims enter the graph; assessments run under the same public standards whoever pays, funders never see or shape a verdict before anyone else, and mandates that attempt to steer conclusions are refused.

Created by extractor · Sep 13, 2026. Every judgment on this page is accompanied by a reasoning trace.