I’m a little confused - I thought their proofs were all driven by Lean proofs - is that not right? So even if the quality of the work is low in some metrics, it either passes the test or not..? No space for changing your mind either way.
> As part of our GitHub repository, we are sharing formalizations of many of the proofs in Lean, a programming language that allows mathematical proofs to be checked by a computer. We will update the repository with more formalizations as we obtain them.
Meaning they published all results before checking all of them, and intended to add more Lean proofs later. In the linked post they state ~42% of the posted results now have formalized proofs, some were added, some verified, and I assume this means that some results turned out to be wrong.
This is extremely disappointing. It means they are sharing unproven work for PR, forcing the mathematicians community to do the verification job for them, while so-called "accelerationists" surf on the hype and help with the pro-AI propaganda.
If your AI tool can help advance mathematical research, share the tool with mathematicians. Using it like this is irresponsible.
"AI will kill us all": no. Greedy humans will kill us all. With AI.
Every option is going to lead to someone shitting on OpenAI for what seems to be a pretty huge accomplishment. There have been opinions written by some mathematicians that OpenAI should just share the work that they have now so that people who are working on any solved problems can know. Which seems reasonable to me.
Lean itself is very hard to get rigorously correct, if you have every tried it yourself. I am not surprised if some AI even tries to benchmaxx Lean 4 by some loopholes
The "irresponsible" part is about how orange buffoons in power will read this, and immediately defund all universities, and mathematicians will lose their jobs, leaving us with nothing but an AI tool that no one can keep in check anymore.
No, it's the exact opposite, actually. They were sharing them early on advice of mathematicians -- they were criticized for being opaque for too long with previous announcements. OpenAI is in ~bad faith, but this isn't a sound criticism.
Also you are deeply confused about what accelerationism is, I believe. Sorry.
Whom? Did they create their own board of mathematicians that would agree with them? See below Terence Tao's blog, sharing a statement from the Association for Human Mathematics.
But that exactly how human mathematicians do things. They upload their research to preprint services like arXiv as they await it to be peer reviewed and accepted into a journal. Why is it okay for mathematicians to publish preprint papers, but when OpenAI does it, it's irresponsible?
Only a subset contain Lean formalizations. And even for that subset, there's the potential that the formalization is semantically off (that is, it's a formalization for a slightly different problem).
If you wrote twenty million lines of Lean to verify something, my suspicion is you've been fuzzing the Lean solver rather than coming up with new math.
How do we even know the premises of the "verified" Lean proofs are correct? The more I think about these results, the more I'm convinced this is like a junior engineer who writes 100 unit tests and shares a screenshot of Pytest being all green, but you check the code and most of them are just doing assert True.
If you go far enough to the edge that's already how math works, since everything is very interpretation-sensitive.
There is however a new problem of scale. Erdös was a human and still managed to create work for an entire generation of mathematicians, how much of a mess will an automathician create?
Sure but if the results have practical implications then the validity will be self evident. If they do not, not much of consequence has been lost.
Interesting that math gets so much attention, when actual advances to material science, biology and chemistry have much higher ramifications and economic benefits. I assume progress there is kept under wraps until they can capture the economic benefits. If they can do that, then the insane valuations may actually be valid.
> I assume progress there is kept under wraps until they can capture the economic benefits. If they can do that, then the insane valuations may actually be valid.
Or progress is not as straight forward in those fields as in math.
I've tried reading one of these, it was an unreadable mess with some strong smells. It might have something to it, but it'd take a decent amount of labor to validate it, especially with how many references to other papers it had.
This is almost certainly the case. I very much doubt that anyone there of any importance in the decision-making process around this actually cares about the math, just the headlines they can get from pushing it out.
I am starting to think do they have some metrics or KPIs that they are trying to fill with this stuff. Pressure to produce anything that at least on first glance sells... Then again they probably are not only place doing that with AI...
I am certainly experiencing what seems like some mania or computer addiction from these technologies. I’ve never been able to produce such results as I can today. I lose sleep staying up late working on it (though to be fair this has always been an issue). But the volume of work is so hard to audit. It makes it difficult to make flawless results. That doesn’t excuse the mode of publication. They could have had humility in their announcement. “We are seeing some interesting results and seeking community validation.” Maybe they did, I did not read their full announcement. But that would have been the right move if they can’t verify something fully.
I’m a little confused - I thought their proofs were all driven by Lean proofs - is that not right? So even if the quality of the work is low in some metrics, it either passes the test or not..? No space for changing your mind either way.
No, in their original blog they wrote:
> As part of our GitHub repository, we are sharing formalizations of many of the proofs in Lean, a programming language that allows mathematical proofs to be checked by a computer. We will update the repository with more formalizations as we obtain them.
Meaning they published all results before checking all of them, and intended to add more Lean proofs later. In the linked post they state ~42% of the posted results now have formalized proofs, some were added, some verified, and I assume this means that some results turned out to be wrong.
This is extremely disappointing. It means they are sharing unproven work for PR, forcing the mathematicians community to do the verification job for them, while so-called "accelerationists" surf on the hype and help with the pro-AI propaganda.
If your AI tool can help advance mathematical research, share the tool with mathematicians. Using it like this is irresponsible.
"AI will kill us all": no. Greedy humans will kill us all. With AI.
Every option is going to lead to someone shitting on OpenAI for what seems to be a pretty huge accomplishment. There have been opinions written by some mathematicians that OpenAI should just share the work that they have now so that people who are working on any solved problems can know. Which seems reasonable to me.
Lean 4 is relatively a new thing, last time I checked the formalization of undergraduate level mathematics isn't entirely done yet.
example https://ai.math.uw.edu/projects/spring-2026/
Lean itself is very hard to get rigorously correct, if you have every tried it yourself. I am not surprised if some AI even tries to benchmaxx Lean 4 by some loopholes
Irresponsible? You can safely ignore the GitHub repo if it bothers you so much. You can safely browse away.
The "irresponsible" part is about how orange buffoons in power will read this, and immediately defund all universities, and mathematicians will lose their jobs, leaving us with nothing but an AI tool that no one can keep in check anymore.
No, it's the exact opposite, actually. They were sharing them early on advice of mathematicians -- they were criticized for being opaque for too long with previous announcements. OpenAI is in ~bad faith, but this isn't a sound criticism.
Also you are deeply confused about what accelerationism is, I believe. Sorry.
> on advice of mathematicians
Whom? Did they create their own board of mathematicians that would agree with them? See below Terence Tao's blog, sharing a statement from the Association for Human Mathematics.
https://terrytao.wordpress.com/2026/10/07/ahm-statement-on-o...
But that exactly how human mathematicians do things. They upload their research to preprint services like arXiv as they await it to be peer reviewed and accepted into a journal. Why is it okay for mathematicians to publish preprint papers, but when OpenAI does it, it's irresponsible?
Only a subset contain Lean formalizations. And even for that subset, there's the potential that the formalization is semantically off (that is, it's a formalization for a slightly different problem).
If you wrote twenty million lines of Lean to verify something, my suspicion is you've been fuzzing the Lean solver rather than coming up with new math.
Well, fine - but my understanding is, if a fuzz-generated Lean proof is correct, that’s end of story. It can’t be “incorrect” if it “passes”.
You might think this is not very useful, maybe - but that’s not a reason to retract..?
If you're fuzzing the solver, you might discover a solver bug.
How? What about the lean verification?
How do we even know the premises of the "verified" Lean proofs are correct? The more I think about these results, the more I'm convinced this is like a junior engineer who writes 100 unit tests and shares a screenshot of Pytest being all green, but you check the code and most of them are just doing assert True.
Were the withdrawn ones actually Lean-checked or not? seems like that matters
Just the other day I was thinking, if unsupervised maths will descend into "oops we found a bug in some code, branch XYZ of maths is no longer true".
If you go far enough to the edge that's already how math works, since everything is very interpretation-sensitive.
There is however a new problem of scale. Erdös was a human and still managed to create work for an entire generation of mathematicians, how much of a mess will an automathician create?
> automathician
I vote for "automathon"
c.f. Italian differential geometry
Sure but if the results have practical implications then the validity will be self evident. If they do not, not much of consequence has been lost.
Interesting that math gets so much attention, when actual advances to material science, biology and chemistry have much higher ramifications and economic benefits. I assume progress there is kept under wraps until they can capture the economic benefits. If they can do that, then the insane valuations may actually be valid.
> I assume progress there is kept under wraps until they can capture the economic benefits. If they can do that, then the insane valuations may actually be valid.
Or progress is not as straight forward in those fields as in math.
Early sentiments are a lot of the write ups still read like slop and it feels very rushed and not very polished.
I wonder if the issue is that not many people within OpenAI can validate the output. Therefore what reads well looks good, and then was published.
If that’s the case in a way its a similar delusion that average people are experiencing with their own AI use.
I've tried reading one of these, it was an unreadable mess with some strong smells. It might have something to it, but it'd take a decent amount of labor to validate it, especially with how many references to other papers it had.
This is almost certainly the case. I very much doubt that anyone there of any importance in the decision-making process around this actually cares about the math, just the headlines they can get from pushing it out.
I am starting to think do they have some metrics or KPIs that they are trying to fill with this stuff. Pressure to produce anything that at least on first glance sells... Then again they probably are not only place doing that with AI...
I am certainly experiencing what seems like some mania or computer addiction from these technologies. I’ve never been able to produce such results as I can today. I lose sleep staying up late working on it (though to be fair this has always been an issue). But the volume of work is so hard to audit. It makes it difficult to make flawless results. That doesn’t excuse the mode of publication. They could have had humility in their announcement. “We are seeing some interesting results and seeking community validation.” Maybe they did, I did not read their full announcement. But that would have been the right move if they can’t verify something fully.
Yes, of course. Wasn't that the whole point? To get more humans involved earlier in the process, with more transparency?