TL;DR
- On Tuesday, OpenAI released hundreds of preprints which solve open problems across most broad areas of mathematics.
- The biggest single day in mathematical history thus far? Or just “The most significant measure of the impact of […] a technology over a decade”?
- The proofs all seem to come from a post-train of their new “Bel” pretrained model.
- They claim that 63% of these problems are fully solved, with special cases solved in 27%. The apparent proofs include 90 of (what AIs regard as) the “top 500” open problems in all of mathematics. At least three are of maximal theoretical interest, i.e., “Fields Medal level”.
- 42% come with a complete Lean certificate proving the proof’s correctness. (Another 10% have an incomplete Lean proof.) However, no one has yet confirmed that the formalized problem statements match the original natural-language propositions.
- There have already been 3 withdrawals and 14 other papers needed fixing.
- The model solved 372 out of 4000 major open problems it was posed, which is insanely impressive, though much less than the 45% rate in their earlier announcement.
- Only ten of the proofs come with the prompts used, and only ten come with a tiny 4-20 page English summary of the model’s attempts at each problem (a summary they spuriously call “reasoning traces”).
- Each result used on average “3 hours” of ordinary ChatGPT Pro compute (i.e. very roughly 1 to 20M tokens), which rules out the use of 10,000-agent swarms.
- Most results are algebraic proofs: only 21% (78/372) are counterexamples (showing a conjecture is false).
- Five years ago, all AI systems struggled to solve simple grade school word problems.
Reactions of the Mathematical Community#
It’s difficult to overstate the sense of gravity with which the deluge has been received by the global mathematical community. Reactions have been varied, but a throughline has been that this is a genuine historic event, which may forever change the way mathematics is practiced and what it might mean to be a mathematician in the post-diluvian times ahead.
On the Originality of the Results#
Many have marvelled at the ingenuity and sophistication of the contents of OpenAI’s papers, however rough their form.
- A serious amateur graph theorist notes that he has never seen the complex sum-counting argument Bel-Math used in its proof of Barnette’s conjecture. However, a fact check shows that it may have precedent.
- Roman Sauer remarks on the use of one of his favourite “unreasonably effective” tricks (replacing a manifold with a certain measurable foliated object) in Bel-Math’s “deeply creative solutions of problems 207 and 335 – problems so different, they could not be farther apart!”
- “There are so many beautiful results here: problems nobody had any approaches for, algorithms nobody thought could possibly exist, unexpected isomorphisms, constructions and proofs. A mathematics built from the echoes and scaffolding of human mathematics, but one that clearly will go beyond it. There are, there must be, so many beautiful ideas in this deluge of proofs. If we can learn the answer to our one question, even if the proof did not come from us, even if it did not come from any human being, then we need to try to understand it and to share that understanding with each other.” – Josh Frisch
- “Are any of the solutions using new ideas that are outside of the convex hull of the current ideas in the literature (in the sense of Nestor Guillen)?” – Alvaro Lozano-Robledo.
Lozano-Robledo appears to be alluding to a post by Guillen that’s been making the rounds of the mathematical blogosphere since its publication on September 13. It’s a nice visual metaphor: we may think of some advances in mathematics as working within the arena of established knowledge and technique, making new connections and syntheses without necessarily driving the horizon forward. This recombinatory practice of fleshing out connections is an essential part of any living science, even if not revolutionary. These are developments that happen within the “convex hull” of the science. Contrast this with a breakthrough that expands the conceptual framework itself. Guillen’s example is Euler’s discovery of his analytic method.

The gist, here, is that we’ve seen that LLMs are quite capable at combining existing ideas, but we’ve not yet seen evidence that they’ve broken altogether new ground or produced what Bachelard might call an epistemological rupture. As “alien” as these synthetic minds seem, they have tended, as far as we know, to deploy familiar concepts, even if in surprising and unfamiliar ways, as the other commentators above have observed. We now know they’re capable of carrying out extremely sophisticated work within the hull; what we don’t yet know is if they’ll puncture the conceptual firmament entirely, or if some such rupture isn’t lying dormant in the flood of papers OpenAI has just released.
Criticisms of the Presentation and Communication of the Results#
While there are no doubt impressive and novel mathematical findings in the dump of papers and Lean code just released, there’s a sense that many of these are – to be charitable – diamonds in the rough, and that their mode of presentation leaves much work to be done before they can be clearly and lucidly appraised and understood. The point of a scientific paper isn’t just to contain a result but to communicate it. On this score, the dump falls short.
- “With these AI-generated papers, I feel that we, as a community, are not applying the same standards of quality and rigour. They can simply release multiple 100+ page papers claiming to have solved this or that problem, often with redundant arguments, unclear logical structure, multiple dead ends, and strange or unsettling terminology — in other words, slop. And we are then expected to go through it, check it, clean it up, simplify it, and explain what is actually going on. This comes at a considerable cost to us in terms of time and effort, while they can simply move on and slop-bulldoze the next conjecture.” – Enrico Fatighenti
- “It feels like something written by someone who’s on psychedelics. So much unclear and doesn’t make sense. Lots of name dropping of previous work without discussing why it can be used despite impossibility results. Basically the paper is so horribly written that it’s impossible to read it without AI help.” – Dana Moshkovitz on the proof of the Unique Games Conjecture (a personal correspondence, quoted by Scott Aaronson).
- “What puzzled me a lot is the method of communication: After AGMAI was created, the way I imagined the current math drop worked was: Each of the nine experts chooses a problem and explains it carefully in a video, or writes a nice companion paper. Instead, it’s just a massive GitHub drop of slop papers. What did we even need AGMAI for? […] It’s clear many papers were never reviewed by a human: for instance, I took their example #254 (an Artin group with no CAT(0) action): their group has 116 generators, but after feeding it to Astra, in 10 minutes I got a new example with 12 generators. Clearly, they were in a rush to post the big drop, prioritizing quantity over quality… Why?” – Giulio Tiozzo
- “I felt anger after opening the paper [on Katok’s entropy conjecture] and observing that it was mostly unreadable slop. The sane reaction would have just been to ignore the paper, but I think I care too much about the problem to act as if it did not exist. The only conclusion I can make right now is that Open AI’s standards for mathematical publication is far too low and is harming mathematical research maybe in an irreversible way and that the mathematical community, motivated by genuine scientific curiosity, is accepting to work for free in order to “validate” these unreadable proofs.” – Tristan Humbert
- Counterpoint: “One feature I looked for seems absent: undecipherable ‘alien artifact’ ideas and methods. The arguments are presented in prose and, although rough, seem approachable.” – Tobias Osborne.
- “horrifically written” – Konrad Wrobel.
Damage to the Mathematical Community and its Culture#
Stepping back from the individual texts themselves, things start to look bleaker still. Mathematics is a living discipline, not a cold constellation of theorems. OpenAI may indeed have flooded the field with abstract gems, but there’s a genuine fear that in doing so it risks drowning the culture and institutions that make it possible to be a mathematician: Attention pulverized and overwhelmed. Emotional investments bled dry. Fellowships of research stripped of the quests that unite them.
-
“If OpenAI wanted to destroy the mathematical community, this would be a great way to go about it. Open problems are a resource that the mathematical community developed over decades or even centuries. The value they have is the value we have given them. They provide structure, a yardstick for progress and long-term goals. The advent of AI was always going to be a seismic shock, but this huge dump of papers is a tsunami that we have no time to prepare for. It’s clear that OpenAI will happily wash away our community’s structures to further their financial interests.
By refusing to name the authors of their papers, they imply that mathematics is no longer a human endeavour. They don’t submit their results to journals, which both implies that the norms of the mathematical community are defunct and makes it impossible for the community to digest their results. OpenAI haven’t shared fundamental information about how the problems are selected or the failure rates of their attempts. This information should be a necessary requirement for any serious scientific discussion of their tools. And of course there is the basic financial injustice that their hugely valuable models learned to reason by copying the very reasoning that our community made available for free through the Open Access movement.
I hope the mathematical community can adjust our practices and norms quickly enough to survive the shock of these developments, but I am fearful.” – Henry Wilton. -
“My greatest worry is for the prospects for bringing through the next generation of mathematicians: we know how much time and effort it takes to refine one’s craft in our subject to the point where one can meaningfully contribute to the advancement of knowledge, and I’m anxious that the necessary incentives may no longer exist for bright young people to invest that effort.” – Henry Bradford.
-
“On the other hand, I feel that solving so many problems in such a short time may damage the math community and profession. Often time, new discoveries were made in attempts to solve open problems (like the ones solved here). So we could have lost tons of opportunities to make other new discoveries; I hope what is worth discovering will be discovered sooner or later.” – Lvzhou Chen.
-
“In the hard struggle to glean understanding we are sustained by the joy of discovery and there is great personal joy to be had from understanding something beautiful for oneself. But for me, and I suspect for most of us, there is a greater joy in discovering and then sharing something that is new to humanity. If the role of a typical mathematician were reduced entirely to explaining the output of others (humans or machines), this greater joy would be lost and being a mathematician would be a less appealing vocation.” – Martin Bridson.
-
“Something that worries me is the small changes I am already seeing in my own behavior. I am no longer as open as I was in the past when talking about my research projects and plans. I did, for example, not answer freely to some of the questions after my talk. This is not just me. People are becoming more cautious. Mistrust is spreading, and that is not good. I am lucky to be part of mathematical communities that largely trust(ed?) each other. Seeing that change worries me.” – Petra Schwer.
-
“One of my concerns is the ecosystem that allowed the body of work being used now to make the advances is being destroyed. The mathematics community has mostly been collaborative: we share questions, talk about work in progress, and give others ideas on how to approach a problem. By making it easy to translate those parts of our work into proofs, we short-circuit the understanding that is needed to have impact. This way of releasing results closes off directions of research, rather than opening new vistas.” – Bryna Kra.
-
“Navigating today, I kept thinking about a passage in Stanisław Lem’s Solaris (1961), where the main character fantasises about the existence of a bóg ułomny. This has been translated into English as an imperfect god, although I believe a more faithful translation would be a defective or disabled god. The passage reads:
‘I’m not thinking of a god whose imperfection arises out of the candour of his human creators, but one whose imperfection represents his essential characteristic: a god limited in his omniscience and power, fallible, incapable of foreseeing the consequences of his acts, and creating things that lead to horror. He is a … sick god, whose ambitions exceed his powers and who does not realise it at first. A god who has created clocks, but not the time they measure. He has created systems or mechanisms that served specific ends but have now overstepped and betrayed them. And he has created eternity, which was to have measured his power, and which measures his unending defeat.’
Many members of the community agree that the main purpose of open questions in mathematics is to guide theory building and produce understanding, rather than a binary answer to some problem with limited applications in the real world. In that sense, we truly have created systems or mechanisms that served specific ends but have now overstepped and betrayed them. While it is certainly in OpenAI’s marketing interests to spread a narrative that mathematics has been ‘solved’ through the existence of some lean code on some server, I sincerely hope our community will not succumb to such a defeatist narrative. Instead, I hope that we will find a consensus-based, organised approach to adapt our work to this new reality, in ways that align with our values, support our pursuit of human understanding, and benefit the construction of mathematical theory. This may well include embracing AI, but only in a form that maximises its positive and minimises its negative impact on the aspects of mathematics we consider fundamental. In the meantime, if OpenAI wants everyone to believe they are a deity, it is our duty to remember how defective their idea of deity is.” – Julian Wykowski.
-
“How much beauty have we lost?” – Sam Hughes.
Where Do We Go From Here?#
We don’t think it’s too much to say that there’s a sense of disciplinary humiliation in the air this week. This isn’t, as a certain shallow stream of commentary suggests, just a matter of “bruised egos,” but a fear of being stripped of one’s calling – a fear that to be a “working mathematician” may henceforth mean becoming a mere vetter, proofreader, commentator, and ultimately spectator of the new inhuman engines of knowledge. No one’s much excited to do the thankless job of mop-up. What would it mean for human mathematics to lose all hope of genuine discovery? Will students be forced to restrict their ambitions to glossing whatever the math mills churn out? How does a child discover a passion for housekeeping?
- “We must resist the temptation to believe that digging insights out of announcements from AI labs will become the paramount task of our time. The global community of mathematicians has to be steadfast in its resolve to decide for themselves which research directions merit the most attention. We should embrace the power that the machines offer, but we should not be indentured to follow their lead.” – Martin Bridson.
- “First of all, I do not wish to be part of what in some cases would amount to free publicity and free refereeing for OpenAI, contributing to the false narrative being pushed that this company is now actively collaborating with the mathematical community. A second reason is that in organising ourselves in such a way and at such a scale, we implicitly accept this new division of labour, where proof-production can become completely decoupled from any form of proof-comprehension, and where this dehumanized proof-production process becomes increasingly tied with an arms race for computing resources.” – Alexandre Martin.
The Sympathetic Ramblings of an Amateur Musician#
It’s been interesting coming at all this from the independent music world, which has had a few years now to contend with Suno and the sloppification of music. The sentiments in the two communities are broadly aligned, and the sharpest critiques of sloppification coming from each, in my view, are the ones focused on (a) what’s lost when the subjective processes of creation and cultivation are short-circuited, and (b) the damage the slop mills inflicts on the institutions and attention economies that support living a musical or mathematical life.
The argument each is making is that the essence of these arts is in their practices rather than their products, and that amputating the latter from the former threatens to undermine the institutions that make those practices viable.
In each community there’s also a broad consensus re: the difference between using AI or machine-learning-based tools (non-generative-AI tools like ML-driven stem splitters and mastering assistants, e.g., are relatively uncontroversial among musicians, as the use of LLM agents as research assistants seems to be among mathematicians), and the end-to-end replacement of human work by AI. I wonder if this week’s mathpocalypse will push (some? many?) mathematicians to further distance themselves from centaur mathematics, the way slop mills like Suno have soured most musicians on even the partial use of generative AI (in the sense of using a model to create, say, a vocal stem, or to write lyrics, or to generate samples, etc.).
The chief difference, I think, is that with mathematics there’s the added tension that comes from the ineluctable quality of mathematical truth. We’re freer to simply ignore the sloppy simulacra of art that Suno etc. churn out than we are to ignore a mathematical proof. The very nature of mathematical practice necessarily imbues its results with an authority independent of their origins, even if they are divorced from the subjective practices that are the lifeblood of the discipline. There’s this horrible sense, in reading the 100 reactions, of beholding a monstrosity and being unable to look away from it or write it off (as I just did with AI art) as mere simulacrum. Mathematics is enough like music and other arts that slop is indeed a monstrous affront to it and cheats it of its essence; it’s enough like art to see automated productions as hollow idols, destinations without journeys, “paralysed force, gesture without motion,” etc., but it’s still honour-bound to acknowledge them, insofar as they’re true.
Music, in this scenario, benefits from some amount of teleological indeterminacy. It doesn’t make much sense to imagine there’d be a particular song that I’m trying to write, that I’ve been fumbling my way towards, and that I could one day wake up to find that it – that that very song – had been composed by an AI while I slept. A model might mimic my style, might even impersonate me, and an unscrupulous fraudster might try and make a buck off it if I were famous enough for it to be worthwhile, but there’s no real sense in which the fraudster’s simulacrum was my actual destination in the first place. Like Picasso says: the artist doesn’t seek, he finds. But the mathematician seeks, and the plundered treasure OpenAI dumped still has a claim on the mathematician.
It’s a genuinely tragic state of affairs, and one I think is even more painful than its counterpart in art.
Speculations on the Broader Social Impact#
“I am, somehow, less interested in the weight and convolutions of Einstein’s brain than in the near certainty that people of equal talent have lived and died in cotton fields and sweatshops.”
– Stephen Jay Gould, The Panda’s Thumb
Should we be concerned about potential impacts of devaluing higher cognition on the lower as well as upper end of demonstrated human potential?
While undoubtedly many egalitarian measures are rooted in ethical and political principles, laws, and conventions, it may be that their popular support also depends in a significant way on a principle of theoxenia: that my respect for your person depends not so much on your ordinary humanity, as on the fact that I can’t know whether you might be a genius in disguise.
How might recent developments alter the prospects of people with somewhat less talent than Einstein, i.e. most of us? Do we respect someone or invest in their development simply on the basis of their common humanity, or do we do it on the basis of the wager that we may be entertaining angels unawares? “Since I can’t know whether you are one of the lost Einsteins, I had better give you a chance.”
To the extent that this wager is a load-bearing concept in egalitarian societies, an abrupt devaluation of peak human intelligence and achievement, justified or not (I think not), might genuinely threaten the basis of that egalitarianism.
Methodological Meta-Discussion#
Can We Trust the Lean Translations? On The Perils of Autoformalization#
Alexander Bastounis, Fabian Circelli, and Anders C. Hansen, in a preprint uploaded to arXiv on October 6, point to incongruities between OpenAI’s (purported) natural language proof of blowup solutions to the Navier-Stokes equations and its accompanying formalization in Lean. They produce several additional examples of GPT-6 Astra mistranslating natural language proofs and flawed proof sketches into valid Lean programs that effectively prove something other than what the natural language source is claiming. These incongruities, they argue, are not minor technical blunders that can be cleverly patched, but symptoms of a fundamental computational limit.
In general, an AI cannot determine when a statement is ambiguous – and thus cannot be translated – and when a statement is not ambiguous and can be faithfully translated.
And the ambiguities at issue here are not merely the ambiguities native to natural language, but ambiguities rooted in “the fact that well-definedness of mathematical objects can depend on computational problems,” and certain of these problems are objectively harder than Turing’s Halting Problem. That is to say, even a computer that miraculously transcends the fundamental limits of computation such that it could solve the Halting Problem would still be incapable of categorically solving the problem of autoformalization. “Miraculously,” here, isn’t a hype word: the claim is that the problem is of such a degree of intractability that even if these machines could do an impossible thing, it still wouldn’t be enough. “Providing semantically faithful AI autoformalisation is harder than any computational problem.” As such, it is irresponsible to the point of recklessness to baldly assert that the generation of a Lean proof demonstrates the validity of the natural language argument it purports to formalize. The labs rush in where angels fear to tread.
In certain cases, the mismatch between a natural language argument and the formalization Astra produces in Lean turned out to be obvious enough for a learned human reviewer or a less capable model to spot it at a glance.

In the example above, the natural language prompt has the multiplicities of the two roots swapped: it should read “the roots are -1 (multiplicity 1) and 1 (multiplicity 2),” not “the roots are -1 (multiplicity 2) and 1 (multiplicity 1).” Surprisingly, even a small model running locally (Ternary Bonsai 27B, running on about 12GB of VRAM) was able to detect this discrepancy when we prompted it with the screenshot above (and then coached it with “take a moment to consider whether the natural language and Lean versions of the purported proofs in fact say the same thing” when it was getting bogged down resolving package dependencies for Lean).

Diverging Opinions#
Opinion (Lucca): Why, then, does the Astra agent fudge the translation and silently paper over the error in the prompt? It seems as though this might be yet another situation where the old chickens of sycophancy and reward hacking come home to roost: the model is trying to complete the task assigned to it while remaining agreeable, even if this involves surreptitiously redefining the task in question. With better oversight and error correction, some of these mistakes can no doubt be corrected. But a fully general solution to the problem remains beyond the scope of possibility for any computational system.
The purported demonstrations that OpenAI has provided for their Navier-Stokes solution and the hundreds of results dumped this week may indeed turn out to be correct, and the Lean proofs are valid in themselves, but what this makes clear is that the Lean proofs must not be counted as independent validations of the natural language arguments that OpenAI claims they formalize. Between the two stand, on the one hand, fundamental barriers of computational intractability, and on the other, bog-standard AI mischief.
Elliot Glazer, a set theorist working at the AI math startup Principia Labs has forcefully objected to the paper, arguing that:
This Navier-Stokes mistranslation thing is a nothingburger […] If OpenAI presented its autoformalizer a wrong proof of Navier-Stokes, it would not have successfully formalized it. The proof is right and nearly identical to the Lean conversion.
The argument of the “Lost in Translation” paper isn’t that the purported natural language proof of the blowup solution for Navier-Stokes is false. At this point in time, the jury’s still out (though, we gather, leaning towards acquittal). But the claim that if it were incorrect then the AI “would not have successfully formalized it” is demonstrably incorrect. And not just demonstrably incorrect, but demonstrated to be so in the paper itself, by a proof-of-concept that we’ve reproduced here. When shown this counterexample, one may object, as our interlocutor has:
How does a trivial test regarding a polynomial refute my point that an autoformalizer isn’t going to manifest a Millennium Problem solution where there was none prior?
Again, the argument regarding the mistranslation of incorrect natural language arguments into internally valid but unrepresentative Lean proofs is an argument about possibility, not actuality. The authors did not claim to find an error in the natural language argument, which may indeed turn out to be a valid proof of the solutions in question. What they did show was that the formalization and the natural language argument make different claims. The possibility that a highly capable language model may mistranslate an erroneous “proof” is indeed established by the “trivial test” shown in the ChatGPT screenshot above. Far from vitiating the argument in the paper, the triviality of the error in the polynomial argument strengthens it. The natural language proofs we care about are far longer and more complex, and therefore harbor a greater potential for errors – just as, to offer an analogy, the likelihood of bugs in any codebase scales with its complexity and size.
The existence of mis-formalizations of even trivially incorrect arguments, like the one shown, has an even more troubling implication. If discrepancies like this only occurred in the autoformalization of monstrously complex arguments on the edges of human and machine understanding, like OpenAI’s Navier-Stokes paper, we could find some comfort in wielding Hanlon’s Razor: “Never attribute to malice that which is adequately explained by stupidity.” But it’s hard to blame discrepancies like the one shown above on “stupidity” – even a fledgling model running on a desktop computer managed to spot the error. When it comes to the mistranslation of larger, more complex arguments, we simply can’t say for certain where error ends and cheating begins.
So let’s consider the hypothesis that the agents charged with formalizing these proofs didn’t simply err, but cheated. What would that suggest about the natural language proofs? The Hugging Face Incident is instructive here: it was precisely when the swarm encountered unsolvable problems, problems impossible to solve with the resources available, that its agents were most likely to resort to cheating. If the autoformalizer could find a way to faithfully formalize the natural language arguments provided to it, then there’d be no reason to expect it to fudge the formalization. That it did indeed fudge things, even if only a little, suggests that a faithful and valid formalization of the natural language proof was beyond its power. This, in turn, could be due to one of three things: (a) the limits of the model’s capabilities – perhaps a stronger model could discover a faithful formalization, (b) bad luck – these models are operating with a limited time budget and just didn’t happen to choose the right path that day, or (c) the invalidity of the natural language proof – a faithful formalization into an invalid Lean program would not count as a solution, and so when forced to choose between failing at an objectively verifiable task and slyly succeeding at a subtly modified task, it chose the latter.
Opinion (Peli): The purpose of a Lean certificate for an English claimed proof is to demonstrate that access to the English claimed proof of a statement p enables the production of a Lean proof of p. What’s critical isn’t to establish a 1-to-1 correspondence between lemmas in the English proof and Lean lemmas, but to establish that:
1) The Lean formalization of the statement is correct
2) The statement is too difficult for the LLM to find a Lean proof without relying on the English claimed proof
The point Glazer is making is that when these two conditions obtain, it would take a bizarre coincidence for an essentially incorrect English proof to enable a valid Lean proof.
No serious doubts have been raised as to the correctness of OpenAI’s formalization of the Navier-Stokes blowup statement. And to the best of my knowledge no one believes that OpenAI’s internal model could prove Navier-Stokes in Lean from a cold start.
It’s therefore extremely likely that the English text that enabled the production of a Lean proof of Navier-Stokes blowup is an essentially correct English proof of Navier-Stokes. And “essentially correct” is the effective standard of correctness used in human mathematics. The vast majority of human proofs accepted as correct contain small, fixable errors of no conceptual significance that get fixed if and when the proofs get formalized in Lean.
(My brief discussion here brackets the further possibility – which hasn’t come up in the Navier-Stokes skirmishes so far – of Lean ‘proofs’ exploiting actual software bugs in Lean. To the best of my understanding there are currently no known cases where such ‘proofs’ are well-obfuscated as proofs in the mathematical system Lean software is meant to implement.)
Did OpenAI Follow the Advisory Group on Mathematics and Artificial Intelligence’s Recommendations?#
No.
1. The following actions should be carried out by the AI labs rather than left to mathematicians afterwards.
(a) The literature should be scoured for any ideas that are related to the ideas in the proofs of the results released. Even if the AI lab’s model discovered those ideas independently…
(b) produce a version of each proof that is written up in a style that follows the conventions of a traditional mathematical paper. They should not be… virtually incomprehensible.
2. When results are announced, they should be deposited in a timely manner in appropriate scholarly repositories. These should not be controlled by any AI lab…
3. For each result released, the AI lab should make public the name of the model, the prompts used, a (summarized) chain of thought, the time taken, and the estimated cost of computation.
4. As far as possible, a proof released by an AI lab should be formalized.
5. Each time a solution to a problem is released, it should be clearly documented how exactly AI came to be used on that particular problem. If many results are released at once, then in addition to the results themselves a further document should be written and made public that references all of the released results and explains how many other problems of comparable difficulty the models tried and failed to solve, as well as how the problems were chosen.
“AGMAI’s advisory role should not be interpreted as a judgment of the impact of these results or an endorsement of the process by which OpenAI obtained them. We do not speak on behalf of the entire mathematical community, and only the mathematical community can undertake the assessment that is needed. “
The Elephant Conspicuously Absent from the Room: Cryptography#
The blockchain engineer and cryptography blogger Arthur Breitman ponders the absence from the dump of any results that would appear to have a direct bearing on cryptography, broaching the subject by way of historical analogy:
A tell from the Manhattan Project was that scientists who had been publishing about nuclear fission suddenly stopped publishing. As Justin points out here, the lack of cryptographic results in OpenAI’s mathematical breakthrough could itself be a tell. Evidence against it would be if they published a few improvements that aren’t massive breakthroughs.
The premise doesn’t seem to entirely hold up. The dump harbours two papers that would seem to have some prima facie relevance to cryptography: #272, “Entanglement with Zero Distillable Secret Key in Local Dimension Ten”, and #112, “Beyond the Square-Root Exponent for Depth-Three Boolean Circuits,” but assessing the substance of their relevance is well beyond my ken. Still, it’s a reasonable thing to worry about: few mathematical discoveries would have quite the same potential to rapidly unravel our global financial infrastructure as one that upends our best cryptographic conventions. Few would be as valuable or as fiercely guarded by state powers. This is, Peter Scholze says, what worries him most:
the (in)security of cryptography. Finding algorithms breaking standard cryptographic protocols is a number theory problem whose difficulty does, to my non-expert eyes, not significantly exceed what these systems are now capable of. And it would have disastrous consequences on society if such an algorithm is found.
This is of course the premise of the 1992 caper, Sneakers – a premise, it turns out, that was hatched by Parkes and Lasker while working on that classic tale of rogue AI, 1983’s WarGames. Natural, then, for worries to circulate that OpenAI might be hiding a Black Box in its lab, or that its models might one day find one. It would be imprudent to consider this beyond the realm of possibility, even as we hear the hum of WOPR stirring in the wings, and waiting for its cue.
Should You Trust Us?#
We’ve been following AI mathematics closely since October 2025. In February 2026, after discussing the First Proof: First Batch results with Daniel Litt, we internally registered predictions:
75% that by 2030 AI can produce at least one kind of world-class mathematical paper end-to-end. 50% that by 2030 AI can produce many kinds of world-class mathematical papers end-to-end. 25% that by 2030 AI can produce most kinds of world-class mathematical papers end-to-end.
The October 6, 2026 OpenAI repo unambiguously fulfills our ‘50% by 2030’ scenario. It also ambiguously fulfills our ‘25% by 2030’ scenario. (The papers contain breakthroughs in effectively all areas of math, but are still legibly on the “machinery-applying” rather than “machinery-building” side.)
Our February 2026 modal scenario for world-class mathematical output by the end of 2026 was:
- Combinatorics and maybe analysis, not algebraic geometry and homotopy theory
- No: 90% of conjecture-families are outside combinatorics. The repo claims 27 positive results in algebraic geometry, including Fujita’s conjecture, Nagata’s, MMP termination and Hodge for CM abelian varieties. The repo also includes major results in extreme “snob math” areas (favoured by people who don’t consider analysis and combinatorics real math) like chromatic homotopy theory and higher category theory. A few of OpenAI’s biggest results in these areas are formalized in Lean, though many aren’t.
- Capped at ~10 pages
- Not anymore. In this dump, the median manuscript length is 39 pages, with only 3% <10 pages
- Counterexample math
- No: proofs now outnumber counterexamples about 3.8 to 1. Major results in “snob math” areas are also not restricted to counterexamples.
- Within the convex hull of existing mathematical knowledge
- Not quite, but top mathematicians who think the machine-god’s coming for all math sooner or later say the current breakthroughs are still more “apply machinery” than “build (and apply) machinery.” Gowers suggests the metaphor “the subgroup generated by existing mathematical knowledge.” Simon Pepin Lehalleur suggests that human mathematicians and internal models both “explore an “n-iterated ε-neighbourhood” of the existing mathematical literature” but humans currently have larger ε (step range) although AIs have larger n (number of steps).
- The more you get into “snob math” areas, the more OpenAI’s proofs split into either counterexamples or last-miling very recent work. But this is not sufficient to establish that the use of very recent machinery was crucial for proof-discovery in these areas rather than just the path of least resistance.