Partially Lean-verified, incompletely open: What OpenAI's math repo actually contains
The 6 October release is unprecedented in its kind: a public GitHub repository, partially Lean-verified proofs, roughly 4,000 attempted problems. But OpenAI has chosen to publish only average compute and no prompts, despite the advisory group at the Institute for Advanced Study having called for full per-result transparency.
What Actually Happened
On 6 October 2026, OpenAI published a public GitHub repository containing 722 machine-generated mathematical manuscripts, organised into 372 research families. The work spans pure mathematics, theoretical computer science and mathematical physics (MSN). Sam Altman has framed the release as the beginning of a "new era of discovery" (The News).
The numbers come from OpenAI itself: the model reportedly attempted roughly 4,000 problems, and each accepted result reportedly cost an average of around three hours of equivalent "ChatGPT Pro thinking" compute (MSN). Coverage uses differing counts — 722 manuscripts, 372 families, and in some outlets "377 results" — and the relationship between these figures is not explained in the available coverage.
The release comes days after a dramatic backstory: a September release in which OpenAI claimed progress on the Navier–Stokes problem had already raised concerns, and according to Gizmodo, the group behind the release is the same inaccessible internal model.
What Lean Verification Actually Guarantees
Part of the proofs are formalised in Lean, a proof system in which a computer checks the argument step by step. Where a proof is fully formalised and Lean accepts it, it has in practice been checked — Scientific American characterises such results as near-certain. But many of the manuscripts still lack formal Lean versions, and OpenAI says that more formalisations will be added as verification work is completed (MSN). The company itself acknowledges that not all results are fully verified, and that some may still contain errors (Times Now/MSN).
The point is therefore twofold: the release is partially machine-checkable, which is far more than most AI releases of this kind can point to — but exactly which of the 722 manuscripts are Lean-verified is not precisely stated in the coverage.
The Headline Results — Unresolved Claims
Among what OpenAI claims are a solution to the four-dimensional Kakeya conjecture, improvements to highly important computer algorithms, and actual progress toward the Riemann hypothesis (Scientific American). The News also reports a claimed partial Birch–Swinnerton-Dyer formula for the leading term — under certain conditions, not in full generality — as well as a zero-free region for a "quasi-Riemann" hypothesis and an upper bound of 9/4 on the matrix multiplication exponent (The News).
This must be read with caution. These are, for now, company claims, not published and peer-reviewed articles. Scientific American estimates that it will take months to work through the results. The News source moreover contains apparent typos (for example "elliptic nerves" instead of what is probably meant to be elliptic curves), which makes the technical details there unreliable without cross-checking. The headline results should therefore be treated as unverified until mathematicians have gone through them.
The Transparency Gap: The AGMAI Recommendations That Were Not Followed
OpenAI states that it consulted an independent advisory group — the Advisory Group on Mathematics and Artificial Intelligence (AGMAI) at the Institute for Advanced Study — before the results were released (Times Now/MSN). But the group has been critical on two points.
First: on 29 September, AGMAI published a blog post in which the group states that "some frontier AI labs test advanced mathematical problems on proprietary models that remain inaccessible to the broader research community," and "we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models" (Gizmodo). The results in the new release come, according to Gizmodo, from the same inaccessible internal model as the Navier–Stokes claim.
Second: AGMAI's recommendations state that a company releasing such results should publish the model, the exact prompt, and the compute time behind each individual result. OpenAI instead chose to publish only average compute time per problem, with some additional statistics — and no prompts. According to Scientific American, the company has stated that it is not bound by the recommendations (Scientific American).
OpenAI spokesperson Lindsay McCallum Rémy nevertheless tells Gizmodo that "AGMAI's advice and public recommendations have informed how we share results. We will continue to incorporate feedback from the community and improve our standards for sharing major scientific advances" (Gizmodo).
The August Meeting and the Unresolved Contradiction
The background is an August meeting between OpenAI and around 40 mathematicians. The company indicated at the time that its models had solved hundreds of long-standing mathematical problems. The participants reacted, according to Northwestern mathematician Bryna Kra, with "a mix of excitement and fear," and the group asked the company not to simply publish them in a blog post or tweet — as had been done with 10 problems earlier that month. She tells WIRED that the input was ignored (WIRED).
Here word stands against word. WIRED reports that company representatives allegedly assured participants that the solutions would not be released all at once — an assurance that OpenAI spokesperson Lindsay McCallum says the company is "not aware of" (WIRED). This contradiction has not been resolved in the coverage, and both sides stand by their versions.
A Community Out of Step With Itself
The response from the mathematics community is not uniformly critical, but divided. Daniel Litt, a mathematician in Toronto, says: "If we want to know the answers to these mathematical questions, I see no reason why we should ask the company to keep them secret from us" (AI Weekly). Terence Tao, on the other hand, has called the pace of AI-generated results from frontier labs "insane" (AI Weekly).
OpenAI also claims that the new, not publicly available model produced almost all of the results in response to a single prompt to a single AI agent — although some results may have required multiple attempts (Scientific American). This is a spokesperson statement that cannot be independently verified, and MIT mathematician Andrew Sutherland urges treating the claim as unverified.
What Remains to Be Clarified
Several things are still open. How many of the 722 manuscripts are actually Lean-verified is not specified. The headline results — Kakeya, quasi-Riemann, Birch–Swinnerton-Dyer — must underpin months of review before they can be considered established. The difference between the counts of 722/372/377 is unexplained in the coverage. And perhaps most importantly: it is unresolved which transparency norms will bind future releases of this kind — AGMAI's recommendations exist, but OpenAI has taken the position that they are not binding.
That makes the release something more than a single news event: it becomes a test case for how frontier labs should share mathematical results — and how large the gap can be between a company's own advisers and its actual practice.

