Two Discrepancies in OpenAI's Navier-Stokes Proof: Text and Lean Code Don't Match

OpenAI claimed this week that an internal model has solved the three-dimensional Navier-Stokes equations – a long-unsolved problem in mathematics.

Illustration: swirling cobalt ink in water fails to pass through a precise metal grid, visualizing a turbulent fluid proof that does not match its formal verification code.
Illustration
Gift article

Two Discrepancies in OpenAI's Navier-Stokes Proof: Text and Lean Code Don't Match

OpenAI claimed this week that an internal model has solved the three-dimensional Navier-Stokes equations – a long-unsolved problem in mathematics. But before the week was out, mathematicians at Cambridge and King's College London documented at least two discrepancies between the proof as stated in natural language and its formalization in the programming language Lean. At the same time, the numbers in the release suggest that the late-September recommendations of the advisory group AGMAI were not fully followed – but that, by AGMAI's own statement, is for the mathematical community to assess. The matter is therefore not settled – it has instead become a test of something larger: how should AI-generated mathematics be verified at all?

What happened

OpenAI announced that the company's internal model has solved the three-dimensional Navier-Stokes equations, a long-standing problem in fluid dynamics, and that it has produced both an analytical proof and a formalization in Lean (cryptobriefing.com, citing Bloomberg Markets). This is, in other words, the company's own claim. As the sources stress: the solution has not yet been officially recognized as a solution to the Clay Mathematics Institute's Millennium Prize problem – a challenge that, according to an opinion piece in Inside Higher Ed, was established roughly 25 years ago.

The Navier-Stokes claim did not come alone. It was part of a massive release of around 722 papers tied to 372 open problems, ranging from algebra and geometry to theoretical computer science, produced largely by an unreleased model (The Conversation). Several papers claim progress on Millennium Prize problems, including Navier-Stokes-related work. Some of the papers have already been retracted or amended. The scale and shape of the release – described variously as "a drop," "a dump," "carpet bombing," and a "mathocalypse" – has, according to The Conversation, sent shockwaves through the mathematical community.

One technical caveat about the numbers: TechCrunch refers to 719 manuscripts in the release, The Conversation to 722 papers. The discrepancy between the sources has not been resolved, so both figures should be read as approximate. The same opinion piece claims that 10,000 internal agents spent 88 hours on the proof (Inside Higher Ed) – but this is an opinion piece, and the details have not been corroborated by other sources and should be treated accordingly.

How verification of AI proofs is supposed to work

To understand why the controversy arose, one must understand the mechanism behind AI-generated mathematical proofs. When AI models solve mathematical problems, they first produce an explanation in "natural language" – the text a human can read and assess. They then attempt to express the result in Lean, a programming language that in theory confirms the proof's correctness by compiling it as code (TechCrunch).

The idea is elegant: natural language is ambiguous, while Lean code can be checked mechanically. The natural-language version serves as the human-readable narrative layer, while the Lean version is the hard guarantee.

But the guarantee only holds if the two versions say the same thing. Lean code can compile perfectly and still prove something other than what the natural-language proof claims – if, for example, the definitions, assumptions, or problem statement diverge between the two. That is why the correspondence between them is the very core of verification.

That is precisely where the Cambridge/King's College paper strikes. It documents at least two discrepancies between the natural-language proof and the Lean code behind the solution OpenAI has delivered to a problem derived from the Navier-Stokes equations, which describe the complex behavior of fluids (TechCrunch).

An important point that is easily lost in the news flow: the discrepancies do not necessarily prove that either version is wrong. A discrepancy can mean that the natural-language version is loose, that the Lean code formalizes a different (perhaps weaker) statement, or that both are correct but not mutually translatable as they stand. What the discrepancies do mean is something entirely concrete: the standards that are supposed to make an AI proof checkable have not been met in this case. There is as yet no single, machine-checkable proof that also corresponds to what is being claimed. The correctness of the Navier-Stokes claim remains an open question.

The standards that were not met

There are in fact recently drafted standards for this kind of release. AGMAI, an advisory group of nine researchers based at the Institute for Advanced Study in Princeton, published guidelines in late September – that is, just before OpenAI's release. Following OpenAI's launch, the group said in a statement that "it is ultimately up to the mathematical community to assess the extent to which our recommendations were followed" (TechCrunch; quote translated from English). It is thus AGMAI's figures and the community's assessment, not the group itself, that underlie the appearance that the guidelines went unfollowed.

The numbers behind the criticism are concrete. Only 10 of the 719 manuscripts included releases of the model's chain of thought (TechCrunch). Without the chain of thought, mathematicians cannot examine how the model arrived at the result – only the end product. And according to TechCrunch, only 42 percent of the released proofs were formalized, that is, translated into Lean or the equivalent – a share that, according to the reporting, falls below AGMAI's recommendations. A majority of the proofs therefore exist only as natural language, without the machine-checkable guarantee.

The combination is problematic: in a range of cases, the community gets neither insight into the model's reasoning nor a formalized proof code. What remains is hundreds of text proofs that must be read and assessed manually – at a pace that does not match the pace of the release. This is the core of the matter: the release created far more mathematics than the community can realistically verify with the tools and information that were provided.

The community's reaction

The criticism has come from several quarters. Terence Tao criticized the release on social media: "Problems are solved autonomously by AI prompters who have no interest in the broader field itself once their original goal is 'solved,' and who do not understand the AI result well enough to answer questions about the result, give talks, or otherwise engage with the rest of the field" (TechCrunch; quote translated from English).

Tao's point is not primarily that the results are false, but that they lack what makes mathematics a living discipline: people who can defend, explain, extend, and debug the work. A proof no one can answer for is, in a sense, not finished – neither socially nor scientifically.

Another, more serious objection concerns provenance. According to the opinion piece in Inside Higher Ed, mathematicians have generally expressed concern about potentially unacceptable and unacknowledged use of recent work by human mathematicians in the AI results (Inside Higher Ed). This should be understood as reported concerns from the community, not as documented facts – but if they prove true, they would raise questions of academic credit in addition to correctness.

What is at stake

OpenAI itself justifies the release by seeking to "enable further progress in mathematics" (The Conversation; quote translated from English). It is this statement that must withstand comparison with reality: a release that includes neither a chain of thought for 709 of 719 manuscripts nor formalization for more than half of the proofs places constraints on exactly the further scholarly progress it claims to be motivated by. Mathematicians cannot build on results they cannot verify, and they cannot verify results without the tools and insight that were withheld.

There is a way out, and it is simple to describe even if demanding: full formalization of the proofs in Lean, so the machine can check them; correspondence between the formalized and natural-language versions, so the discrepancies the Cambridge/King's College paper has documented are closed; and release of chains of thought where the community requests them. AGMAI itself points to where the decision lies: with the mathematical community.

Until that assessment is made, the status is as follows: OpenAI has claimed to have solved a Millennium Prize problem and to have published hundreds of results. The Clay Institute has not recognized the solution. An input from mathematicians at Cambridge and King's College London documents discrepancies between the proof's two versions. The advisory group's recommendations appear not to have been followed. This is not a verified breakthrough – not yet. It is a claim under test, and the testing has barely begun.

All the sources in this story are secondary reporting; OpenAI's own announcement, the released papers, and the Cambridge/King's College paper should be read in the original to verify the details before any final conclusion is drawn.

AIMag.no
AIMag.no
The AIMag.no editorial team covers artificial intelligence, tools, research, and regulation.

Get the best of AI MAG in your inbox

News, analysis, and ideas at the intersection of AI and society.