DeepMind Claims Machine-Checked Proofs for Nine Erdős Problems — Two Open for 56 Years
Google DeepMind's mathematics agents are said to have produced machine-checked proofs for a range of previously unsolved problems — including nine Erdős problems, two of which had resisted mathematicians for 56 years. But for now the figures come from a single secondary source, and it is worth looking closely at what has actually been confirmed.
What is claimed to have happened
According to a post from Cryptobriefing on October 2, 2026, which reports results from DeepMind's research, Google's Gemini-powered math agents have found proofs for a series of unsolved mathematical problems. The results are said to be described in a detailed arXiv article published sometime around mid-2026.
The figures given are concrete — but all should be read as DeepMind's own claims, relayed through Cryptobriefing:
- AlphaProof Nexus is said to have solved 9 of 353 previously unsolved Erdős problems. Two of the nine had reportedly stood open for 56 years. The same system is also said to have proved 44 of 492 conjectures in OEIS, The On-Line Encyclopedia of Integer Sequences.
- Aletheia, another system, is said to have evaluated around 700 open Erdős problems since its rollout in December 2025, finding or constructing solutions for 13 of them — four of the 13 are described as new, autonomous solutions.
It is worth noting an unresolved detail right away: Cryptobriefing's own headline refers to "five unsolved mathematical problems," but that figure does not reconcile with the larger tallies of nine and 13 in the article itself. The basis for the "five" figure is unclear, and we have not adopted it as the story's anchor.
How the system works
What separates this kind of system from a language model writing plausibly looking mathematical texts is the verification mechanism. According to Cryptobriefing's description of DeepMind's framework, AlphaProof Nexus runs multiple independent prover subagents in multi-turn reasoning loops, powered by Gemini 3.1 Pro. The subagents thus work separately on the same problem, and their results are then graded by automated formal proof checkers.
This is the point that makes the results more than just text: a formal proof is an object that a machine can mechanically check step by step, much as a compiler checks code. If a proof passes the checker, it is verifiable regardless of whether the idea came from a model. The model may have good and bad ideas — but an approved proof is either correctly formalized or it is not, and that is not determined by the model's own confidence.
Aletheia is built on Gemini Deep Think, the same model family behind Google's performance at the International Mathematical Olympiad in 2025: there the system scored 35 of 42 points, corresponding to gold-medal level, and solved 5 of 6 problems within the same time limits as human participants. The competition result gives an indication of the development path these proof agents build on.
Verification — and what humans still have to do
According to Cryptobriefing's account of the research, the results were reviewed by domain experts who confirmed both the proofs and that the formalization of the original conjectures was faithful. The latter point is not trivial: translating a solved conjecture into formal mathematics can itself introduce errors, and a proof of something other than what was actually asked is not worth much.
At the same time, the research acknowledges, per the source, an important limitation: current AI systems still require human oversight to assess novelty — that is, humans must confirm whether a result is actually new. In other words, the machine can prove something correctly, but it is still humans who decide whether the proof is a genuinely new solution to an open problem or a restatement of something already known.
The cost side is also worth noting: the computational cost per problem is given as a few hundred dollars. If the figure holds, it is a dramatically low price measured against the years of expert work some of these problems have represented — but here too, the figure is so far known only through the same secondary source.
What this could mean, according to the people building the systems
DeepMind researchers, among them Pushmeet Kohli, have per Cryptobriefing emphasized that the agent-based approach could extend beyond pure mathematics, with fields such as combinatorics, quantum optics and algebraic geometry named as possible target areas. These are the company's own statements about future applications, not documented results — and should be read as a roadmap claim rather than a summary of fact.
In any case, it is the mechanical verification that makes the prospects interesting. If agent-driven proof environments can explore large parts of a problem space at a fraction of the cost, the economics of certain parts of mathematical research change — not by replacing mathematicians, but by allowing the systems to serve as a broad, cheap triage process before human expert time is invested. The research's own caveat about novelty assessment, however, shows where the boundary currently lies.
What we cannot verify
So far, everything above rests on a single source. Several factors limit how firm the story's footing is:
- No primary source is available to us: neither DeepMind's own statements, the reported arXiv article, nor the documentation for AlphaProof Nexus and Aletheia. All figures — 9 of 353, 44 of 492, 13 solutions, a few hundred dollars per problem — rest on a single secondary account.
- The arXiv article is not precisely identified: the source gives only "around mid-2026," with no title, authors or exact date. The expert review therefore cannot be traced from here.
- Cryptobriefing is a cryptocurrency-focused outlet, not a mathematics or AI research publication, and the reliability of its summary of research results is unclear.
- It is unclear what "unsolved" means exactly — whether the problems were open when the AI proofs were found and independently confirmed as new, or whether some had already been solved by humans. The fact that the research itself points to the need for human novelty assessment suggests this question has not been fully answered.
- Which specific Erdős problems were solved is not stated anywhere in the source, so nothing can be said about that here.
The conclusion is therefore split: the technical mechanism — multiple independent prover subagents with formal proof checking — is well described and sound in principle, and the figures are specific enough to be verifiable if the arXiv article exists as described. But until the primary documentation is available and independent mathematicians have confirmed the results, this is DeepMind's claims relayed through one secondary source — not a documented revolution in mathematics.

