News · 2026-10-08
OpenAI withdraws three math manuscripts as its mass release enters review
OpenAI withdrew three related mathematical manuscripts and revised fourteen others on October 7, bringing its public collection to 719 manuscripts across 372 result families. The company now reports 300 top-line results formalized in Lean, a proof-checking system. The corrections make the release an evolving research record whose publication count, formal-checking status and human review must be tracked separately.
Key facts
- OpenAI’s updated collection contains 719 manuscripts, including 300 formalized top-line results.
- The October 7 history records three withdrawals, fourteen revisions and six additional formalizations.
- The work comes from OpenAI’s unreleased internal model, not a newly downloadable checkpoint.
- The primary correction record is OpenAI’s repository history.
The meaningful event is the beginning of an unusually large review process. OpenAI’s October 6 announcement, titled “Sharing AI progress in mathematics”, put hundreds of manuscripts into the public record together. A result family can include a main argument, consequences and alternate proofs. Counting files therefore does not tell readers how many distinct open problems have been settled.
The first corrections have a traceable dependency chain. OpenAI says a sign error invalidated a cancellation argument in “Algebraicity of Weil classes on split abelian eightfolds.” Two other manuscripts depended on the affected construction: “Algebraicity of Kuga–Satake Correspondences for K3 Surfaces” and “The rational Hodge conjecture for products of K3 surfaces.” All three were withdrawn. That is stronger evidence than a social-media allegation because the company’s own history names the problem and its downstream consequences.
The repository README describes roughly 4,000 problems posed to the internal model and an average of about three hours of ChatGPT Pro thinking compute per result. Outputs were organized and filtered into families and manuscripts; some build on earlier outputs. The collection is therefore a curated portfolio rather than 719 unrelated experiments. The published material supplies no collection-wide count of human-refereed results.
Formal checking is valuable, but its scope needs precision. Think of Lean as an exceptionally strict accountant checking a completed ledger. It can establish that the entries balance under the rules it was given. It cannot independently certify that a transaction was described correctly before it entered the ledger. In mathematics, that means checking the encoded theorem and proof differs from checking that they faithfully express the intended problem and accompanying prose.
Alexander Bastounis, Fabian Circelli and Anders C. Hansen make that distinction concrete in their Navier–Stokes correspondence paper. One comparison finds that a cited formal estimate needs five additional derivatives where the written estimate needs four. A pressure-flux comparison finds a different dependence and proof route. Those are specific objections to the relationship between artifacts. The paper does not say Lean accepted an invalid proof, does not declare the final theorem false and expressly avoids a verdict on the written proof’s correctness.
The reception also contains several distinct positions. Scott Aaronson’s primary post expresses excitement about potentially field-defining mathematics while describing the difficult human work of understanding the manuscripts. Aaronson wrote that “no human has understood just about any” of the proofs at the time of his post. His reported exchanges with Dana Moshkovitz reflect provisional engagement with the Unique Games Conjecture argument, not a completed referee report. In the Hacker News discussion, readers debate whether a correct formal theorem can establish the result even when the prose takes a different route.
Institutional objections concern process as well as correctness. The Association for Human Mathematics statement urges mathematicians to discontinue work with OpenAI. The separate advisory group’s statement says its involvement does not endorse the release or assess its results. Those positions should not be merged, and a guest statement hosted on a famous mathematician’s blog should not become that mathematician’s personal manifesto.
For readers, the practical next step is following individual claims through corrections, statement checks, independent understanding and useful exposition. The existing lessons on proof assistants and autoformalization explain why these checks complement each other. Publication volume can indicate a large new workload without establishing that every claimed discovery has survived it.
The honest caveat runs both ways. Three withdrawals do not invalidate the remaining collection. Conversely, 300 formalized top-line results do not imply 300 externally refereed manuscripts or faithful formal coverage of every lemma. The strongest claim today is that OpenAI’s mass release has entered a visible correction and scrutiny cycle; the final mathematical significance still depends on result-by-result assessment.
Key questions
How many OpenAI math results are now formalized in Lean?
Did the Navier–Stokes critique disprove OpenAI’s theorem?
Did OpenAI solve 90 of mathematics’ 500 most important problems?
Cite this
APA
Ground Truth. (2026, October 8). OpenAI withdraws three math manuscripts as its mass release enters review. Ground Truth. https://groundtruth.day/news/openai-math-release-withdraws-three-manuscripts.html
BibTeX
@misc{groundtruth:openai-math-release-withdraws-three-manuscripts,
title = {OpenAI withdraws three math manuscripts as its mass release enters review},
author = {{Ground Truth}},
year = {2026},
month = {oct},
url = {https://groundtruth.day/news/openai-math-release-withdraws-three-manuscripts.html}
}
Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.