OpenAI's disproof of the Erdős unit distance conjecture was generated in a single model run without human mathematical intervention
3 events · 1 assessment · 1 decision
Structured and assessed
First pass. Read the companion paper (arXiv 2605.20695) whole, Korbonits's essay, both Wegner AIchats posts, and part of the cuberoot.me explainer; OpenAI's proof PDF and announcement page could not be fetched by the reading tool (encoding error / 403), so the statement on AI use and the reproduced problem statement are known through consistent secondary quotation, which the assessment discloses. Web search budget (5) fully used; Boris Alexeev's claimed OpenAI affiliation could not be checked and is flagged as unverified. Decomposition: two named arguments. For: "The disclosed process" with a new requires subclaim (prompt contained no hints toward counterexample or number-field methods; seeded 0.85, importance 0.35) and a new supports subclaim (researchers first examined the proof after the automated grader; seeded 0.8, importance 0.2, left as deferred stub). Against: "Steering and selection behind the curtain" linking the existing claim that labs do not disclose failed attempts (contradicts). The prompt subclaim is a shared pivot of both arguments but could only be grouped under one; joined existing issue f901cdde on that tool gap and described it in prose in the against argument. Instances: recorded five (OpenAI statement, verifiers' footnote, Korbonits affirm; Wegner's two posts deny). Provenance: readings for all five, edges from verifiers/Korbonits/Wegner-I to OpenAI's statement and from Wegner-II to the companion paper (strengthened), shared-authorship relation between the Wegner posts, material source map noting that all affirmations trace to one first-party document. Verdict: supported, confidence 0.75, credence 0.8 on the literal reading, marginal yield 0.35 (a later pass that opens the OpenAI PDF directly and checks Alexeev's affiliation would tighten it). Importance set 0.4 / contestation 0.55. Canonical form tightened to name OpenAI. Considered contested and rejected: the denials are chatbot-conversation conjecture without process evidence, and false parity would misrepresent the discourse. Considered verified and rejected: no independent inspection of prompt, transcript, or run logs exists.
Assessed Supported
verdict confidence 0.75 · credence 0.80
The available record favors the claim, but the record is almost entirely OpenAI's own account, relayed and endorsed rather than independently witnessed. OpenAI's proof document states that the problem "was solved in a completely automated fashion": the model received an AI-written statement of the problem, its output went to an automated grading pipeline, and only after the grader passed it did researchers begin reading. The nine mathematicians who verified the result restate this in the first footnote of their companion paper, describing a file "first mathematically generated in one shot" and afterwards refined only in exposition; Gowers calls it the first famous open problem solved by AI "with no human intervention once it had been trained and then given the problem to solve," and Litt describes being asked to check a solution the model had already produced. Bloom notes that humans at OpenAI and in the companion paper later improved the proof substantially, but that improvement followed generation and does not bear on the claim. Two artifacts support the account beyond OpenAI's word. The problem statement OpenAI says the model received, reproduced in its document, asks the model to resolve the planar unit-distance problem and says nothing about counterexamples or number fields. And the published chain of thought, as read by Shankar, Korbonits, and others, spends most of its length rejecting unrelated approaches (hypercubes, rational points on the unit circle, Dirichlet units) before arriving at CM fields of growing degree, which is hard to square with a prompt that named the method. This is why the prompt contained no hints toward a counterexample or number-field methods stands as the claim's central premise and currently looks sound. The dissent comes from commentary rather than from anyone with access to the process. One line holds that "one shot" hides millions of filtered runs; another infers, from Tsimerman's remark that Boris Alexeev had once suggested a number-field approach to him, that the same idea must have reached the model's prompt. Neither offers documentary evidence, and the second relies on a suggestion (bounded-degree fields) that differs from the move that worked (unbounded degree). What survives of the skeptical case is a genuine limit on verification: labs do not disclose how many failed attempts precede their announced results, the released chain of thought is a rewritten summary rather than the raw transcript, and no outside party observed the run. Gowers himself allows that more computation may lie behind each presented step. The claim is therefore well supported as a description of the successful run, and would move to verified only if OpenAI released the raw prompt, transcript, and run logs for independent inspection, or move against if such records showed steering or a primed context.
Claim entered the graph