Buried under the drama is the fact that OpenAI is claiming that an internal model they’ve been training for less than two weeks is more than twice as capable in mathematics as Astra, which was only made public a week ago. Even if this improvement is limited to mathematics, that is an astounding feat.
show comments
hdivider
My take:
1. It shows what even this wave of AI can actually do.
2. I wish it were done by different folks, ideally under some kind of public control like NASA research or the NPR model.
3. Keep in mind: natural science is different. It's not always a matter of computation. Computer science folks often struggle with this -- but this virtual world here does not actually exist. Everything is physical, including information. Any natural science PhD or otherwise knows just how complicated nature actually is -- e.g. mention any research topic and try to encapsulate all the relevant phenomena present there. Pure mathematics is different because we define the problem, rarher than explore nature. We are in my view far away from removing humans in natural science R&D. Advancements in AI however can greatly assist us in all natural sciences, which is already beginning to happen.
show comments
tiborsaas
> We’re sharing a solution to the Navier–Stokes existence and smoothness problem, one of the Millennium Prize Problems. This proof, produced by an internal OpenAI system, shows that the dynamics of the Navier-Stokes equations for fluid motion can develop a singularity in finite time. We’re sharing both a writeup of the proof and a formalization in Lean.
WOW?
show comments
mewse-hn
"we cannot rule out that de-identified data derived from their usage of our products helped improve our models ."
What a landmine sentence to bury in this report, you can't rule out your models were spying on other researchers?
Unlike the vanilla read of the OpenAI press release, it is much more unfiltered and outlines some particularly aggressive behavior by specific OpenAI employees
show comments
recitedropper
Sad turn of events for our world. After watching the behavior of the most senior OpenAI researchers on twitter, I feel even less confident in them as a team to be shepherding this much capital and compute.
The dark forest awaits..
show comments
keeda
It's low-key funny that OpenAI attempted the problem because they thought somebody else had already solved it, but turned it had NOT in fact been solved!
It's also unfortunate that such a potentially momentous occasion is overshadowed by so much drama. Which I suppose is expected given the technology and the people involved are so polarizing.
closetheloopdev
From my reading of the announcement:
- There are at least two versions of a model more powerful than Astra at OpenAI at the moment.
- The less capable version was used to solve the unforced Euler problem (while the one solved by Levent Alpöge and Tristan Buckmaster was forced Euler) with 100 agents.
- The more improved version was used to solve Navier-Stokes, given the results of the unforced Euler problem from their earlier attempt, with 10000 agents.
- OpenAI initially tried a shotgun approach against the 6 Millennium Prize Problems until it emerged that Navier-Stokes was the most likely to succeed.
So the timeline was:
Shotgunning 6 open Millennium Prize Problems -> solved unforced Euler problem with 100 agents -> concentrating on Navier-Stokes with 10000 agents -> solution.
If so, that is fantastic development and a huge success (despite all the drama surrounding it)! Congratulations!
show comments
intenex
I think this is clear evidence that AI models are now at the far frontier of mathematics innovation and discovery and exceed human limits.
This specific problem having had a $1 million bounty on its head and still remaining unsolved for 26 years after the bounty was placed is pretty clear evidence that many of the world's best human mathematicians would have solved this problem if they could have, and none were able to until LLMs came along.
Hard to claim at this point that LLMs aren't capable of novel STEM creativity and genius to a degree that will soon far surpass that of humans.
If anyone has counterpoints to this I'd love to hear them!
show comments
piker
"... The point remains that there is a substantial opportunity cost in converting a historically productive and motivating problem (such as Navier-Stokes regularity) into a mere viral social media post advertising some benchmark progress, rather than actually advancing the field and developing the next generation of both problems to ask, and people to work on them."
This really leaves a bitter taste....
"On Tuesday, September 1, we heard rumors that two Millennium Prize problems had been resolved. Inspired by these rumors and by the step change in performance of our internal model, we launched an effort to evaluate it on all open Millennium Prize problems and a few other high-impact problems."
IPO+rumour driven research.
I appreciate the achievement, but it doesn't feel right.
show comments
highfrequency
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models
This is the crux of it. If Tristan's work and insights were not used to train OpenAI models, then this just looks like a case of hyper-competitive academic sniping that has been going on for decades (check out Watson and Crick!) accelerated by AI as a tool.
But there is one huge question: did Tristan opt out of model training for his ChatGPT and Codex sessions? If the answer is no, then this seems fair game. If the answer is yes, then OpenAI's ambiguity is strongly suggestive that opting out does not mean what they imply it means.
show comments
railgunmerlin
Does seem like they gloss over Alpöge and Buckmaster's work with the following
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .
Which seems a bit irresponsible/rash?
show comments
Jonasori
the context here is super important, for those who haven't seen it yet. OAI maybe just trained on a real researchers solution and then celebrated having scored the goal unassisted save for the brief commentary at the bottom of this blog post. Here's the other side.
In chess, a grandmaster just needs to know at what moment in a game there's a critical move to gain a significant advantage over their opponent. They don't need to know the move itself.
OpenAI got wind that a millenium problem was being solved. And that feels a bit like the critical move in chess. That is - it was a signal that AI advanced far enough that it would be worth spending a lot of time and resources solving a millenium problem.
show comments
lanthissa
5 million messages, 300b output tokens, done in 5 days, and achieving something humans couldn't.
the first "Country of geniuses in a datacenter" moment.
show comments
webcoon
"Across all attempted problems, the agents sent 4.9 million messages and used about 300 billion output tokens."
At a conservative estimate of GPT 6 Astra pricing, this would have cost upwards of 15 Million dollars for anyone using the OpenAI API!
To me this is the one silver lining. Yes, they can solve millennium prize problems, but it still costs a fortune.
pu_pe
OpenAI thinks of this as a scoop, and it is, but the possibility that they trained the model on the prompts of the other mathematicians they were competing with will leave a terrible taste on every scientist's mouth. Seems like yet another advantage of using open models right here.
show comments
ccppurcell
Reading between the lines here, and taking an admittedly very negative view of openai, but they train on user prompts. So if they hear a rumour that someone is about to make a big breakthrough, they have an incentive to scoop by running the model and hoping the solution is in the new training data. Also the statement from the mathematicians in question alleges that they tried to pressure him into academic malpractice. Just appalling timeline we're in, cheers.
aizk
People had joked a couple years ago "Well if they solve a Millenium problem it's AGI"... Well here we are.
show comments
rfgplk
Something I've been going on and on about for months now and no one seems to listen. LLMs today are allowing _anyone_ to access cross-discipline knowledge that was previously entirely inaccessible without a) extremely deep pockets or b) a massively talented and varied team. In fact, contrary to what the masses seem to think LLMs are actually _better_ at hard cutting edge physics/math problems than they are at frontend web stuff (paradoxically). This is why I'm advising most people to start pivoting into much harder to penetrate domains (historically hardware, aerospace, robotics, biotech). Most fields are in their infancy (see the sad state of embedded development) and the gains to be had are massive.
show comments
Reubend
It's great that important discoveries like this can now routinely be accompanies by formalized proofs. The fact that it's being released alongside a Lean proof from Day 1, rather than the Lean proof being released months or years later, is super helpful for verifying that it's correct.
show comments
minimaxir
> Across all attempted problems, the agents sent 4.9 million messages and used about 300 billion output tokens
Don't even try to do the math on how much that would cost at normal API prices. And we don't even know how much more expensive this internal-only model would be!
show comments
cv5005
Maybe a naive question, but how does one know that a particular lean proof is actually a proof of what one thinks?
Like, ok the logic checks out and it proves something, but there's still the problem of does this logical result actually prove the initial question that was asked?
show comments
thomascountz
At all times we maintained the same strict safeguards that we apply to all our frontier model evaluations, including monitoring and isolation.
Maybe just don't mention that bit, OpenAI.
show comments
hypersoar
I dropped out of a math Ph.D. in 2018, and I'm increasingly glad that I'm not in math research, anymore. While it's cool that we can get these results, I don't think that I'd enjoy being a post-AI mathematician.
lwansbrough
It would be nice if one of these models would produce a novel theory or advance the field in a positive direction.
Most (all?) of the big discoveries have been counterexamples, which is just sort of a systematic tearing down human ingenuity. I know that counterexamples are an important part of progress and discovery, but it just feels bad to me.
But I'm not a mathematician, maybe I'm totally misreading the vibe.
show comments
alasano
I don't know about you guys, but I'm hyped about the future.
Cure all illnesses Utopia or Robot Wars Dystopia, both are pretty exciting.
show comments
matteoraso
This is undeniably epochal, but I can't help but notice that this is yet another example of AI disproving rather than proving something. Is this just a coincidence, or does AI slightly struggle with proving theorems?[0]
[0] Struggle relative to its ability to disprove, not struggle relative to people's ability to prove theorems.
show comments
hexomancer
> On Tuesday, September 1, we heard rumors that two Millennium Prize problems had been resolved
What's the other one?
show comments
modeless
So the timeline is:
Aug 28: OpenAI starts training a new model.
Sep 1: OpenAI sees a rumor on Twitter that two Millenium Prize problems were solved and starts their own effort to attack all the prize problems using the new (4 day old!) model.
Sep 3: The new model makes some progress toward Navier-Stokes. Based on this progress, OpenAI focuses on Navier-Stokes over the other Millenium Prize problems, using several approaches in parallel.
Sep 5: Navier-Stokes is solved. Assuming Astra API prices, $15m in output tokens were used by the whole effort.
In this account of the story, no specific information about Tristan and Levent's work is used to inform OpenAI's approach. The focus on Navier-Stokes and the choice of approaches to pursue came from OpenAI's own progress, not specific knowledge of Tristan's concurrent work.
There is a caveat that they "can't rule out" the possibility that Tristan's Codex data could have been part of the training set of the new model, though it is described as "unlikely" and the proofs are substantially different.
This timeline is insane. Navier-Stokes was solved start-to-finish in 5 days? A model in training for at most eight days dramatically outperforms Astra and Fable, and not just in mathematics?
show comments
coffeeaddict1
This has to be one of the most important moments in the history of mathematics. We now have a non-human intelligence capable of solving one of the most difficult problems in mathematics.
3m4r
This is a great day to re-read Ken Thompson's "Reflections on Trusting Trust":
>To what extent should one trust a statement that a program is free of Trojan
horses? Perhaps it is more important to trust the people who wrote the
software.
The modus operandi is now for the AI companies to watch if someone does something in the open like Kevin Buzzard on FLT, use their research and scoop them with brute force.
Or, in this case, stealing prompts from competitors.
Do not use stealing chatbots for research even if you think you have data agreements. The people running these companies have worked on hookup apps for Christ's sake. Get real.
boardwaalk
I do mean to be critical here. I wish there was better moderation so I could find more conversation about the actual discovery here. There are multiple threads on this and I keep scrolling and only seeing more conversation about the drama. Which is about the least interesting thing IMO. I suppose I’m whistling in the wind here and not helping the situation, but damn.
itvision
There's something sinister or crazy good in the article.
OpenAI already has a model that is at the very least twice as smart as Astra.
Oh god.
show comments
fittingopposite
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .
Wow. That sounds like an admission of guilt.
Chinjut
What is going to become of life for those of us who do not work at AI labs and are unlikely to be hired by AI labs, despite all the years we put into learning math, coding, etc? Those of us who made the mistake of studying anything other than machine learning. How will we make a living? (We don't live in a world that seems likely to distribute gains widely instead of largely to the handful of already mega-rich.)
show comments
vatsachak
Called it. AI wins a fields medal before managing a McDonald's
nadermx
I've been working on this problem for what seems like for ever. Kudos to the OpenAI team.
For those of you who don't care about the drama and want to see this distilled to 3 lines:
So what took an autonomous agentic system using a significantly more powerful internal model, totaling multi-millions of dollars of compute in training and inference, was likely to already be solved by a team of a few humans with an orders of magnitude smaller LLM budget, had OpenAI not been foaming at the mouth to jump the shark and claim “AI solves Millenium Problem.”
Also it sounds like the human research effort spanned weeks if not years from Tristan’s statement so it is extremely likely the work and prompts of these human researchers was used in the OpenAI knock-off.
>so far the proof looks more along the lines of another euler blowup proof we had, off of whose ansatz naming we were making really stupid puns like “smooth criminale”, unlike the much better “ideal fluids explode”, Tristan
There are too many ambiguities around OpenAI. Unanswered questions making this ambiguity more.
Why they didn't properly explain to Tristan about usage of their data.
If I run Codex against a project that includes a private API key, is there a chance a future user of ChatGPT could ask for an API key and get back mine?
I've actually asked someone at OpenAI this question and they said that was the "regurgitation" problem and is something which they actively work to prevent happening.
That's reassuring, but I want to know more. I still don't have an intuitive understanding of what kind of data I should avoid sharing with a model if I'm worried about that data causing me problems when it's used for future training.
Is it safe for me to brainstorm future directions for my company with a model, or might that risk someone getting that information in response to a prompt like "What potential directions could company X consider in the future?" in six months time?
show comments
btilly
The problem that I want to see them tackle is formalizing the classification of finite simple groups.
Everyone uses the classification. Nobody has great confidence in the proof. Nobody understands it. There are attempts to reprove it.
If it can be formalized, that would demonstrate that AI is ready to formmalize all of mathematics.
bhouston
What happens to real fluid in this particular cases?
If the singularity is in the physical space?
Is this just a result of ignoring things like friction and energy dissipation via heat, etc?
show comments
harhargange
I have a dumb feeling that the proof will be wrong with serious flaws but that will be found out only after the ipo
twobitshifter
>The groups varied in size, and the group that produced the Navier–Stokes resolution involved on the order of 10,000 concurrent agents… The agents arrived at their resolution on Saturday, September 5, about 88 hours after the first agents were launched.
The Millenium Prize is $1M, what is the ROI? (Edit: since I was not clear, and confused some - I mean for a hypothetical of a third party paying commercial rates to use AI to solve mathematical challenges and claim prize money, not for scientific value alone or as a promotion of an AI lab’s capabilities)
My napkin math - If you get 33 output tok/s each agent will burn 10.5M tokens over 88 days. At $50/MTok (Astra cost), that is $525 per agent. With 10,000 agents, you’d spend $5,250,000 to get back a million.
(We also know that they were running more groups that varied in size and this model is a generation ahead of astra)
show comments
mapmeld
> Our goal in releasing this result is to report on the substantial progress of our AI models. We do not intend to claim the Millennium Prize for this result.
Does OpenAI have a policy of not claiming math prizes like this, or is this them trying to avoid any concerns (right or wrong, I'm sure we will hear more in the future) about how they got there?
show comments
Kotlopou
For now I think more or less the same thing as with all recent math announcements: This is in a range where human work still exists (see Terry Tao, (1)). I wonder whether the trend will extend into the problems that (as far as I can tell) are considered complete brick walls right now -- P vs. NP, Collatz, Goldbach, odd perfect numbers, problems that aren't part of any research program. (2) In other words, is the progress coming from putting together vast amounts of existing work and computational power, or is it more from RLVR and self-play and autonomous effort?
The answer to this will obviously shape the near future of mathematics, but there's also something even bigger than that at play: It has always been the case that the questions in math were stronger than the answers; you have stuff like Fermat's great theorem that is easy to state but monstrous to prove. This seems to be a property of mathematics, not of humans... but is it true?
A question by Scott Aaronson from 2011 (3) about P vs. NP seems relevant here: "Will humans manage to prove P≠NP before they either kill themselves out or are transcended by superintelligent cyborgs? And if the latter, will the cyborgs be able to prove P≠NP?" Later, he notes that if P≠NP, "once the robots do overtake us, they won’t have a general-purpose way to automate mathematical discovery any more than we do today".
Elsewhere in the thread, others have calculated $15mm at API rates for just the output token. (So I’ll assume this cost about that much, taking input and human researcher time.)
I wonder whether a team of 60 mathematicians working solely on this for a year would have cracked this. (Assuming $250k total compensation.)
show comments
dataflow
Is blockchain going to finally be the solution to something?
I'm only half joking. Should researchers perhaps put hashes of their attempts on a public blockchain tied to their own public keys, verify their claims asynchronously, and then whoever reveals the first believable attempt gets the credit?
I know some people started doing this years ago but now it might need to become standard practice.
> [T]he group that produced the Navier–Stokes resolution involved on the order of 10,000 concurrent agents.
amberjack
Seriously starting to think we are not going to make it out alive of the near-future.
ninjahawk1
The problem is the precedent this creates. For non-famous people using public APIs like this it could mean AI companies sucking up the information and throwing millions in compute at it.
The sequence for Navier-Stokes was that these researcher spent a year working on it, then they published a possible breakthrough, OpenAI then spends $15M within a couple days to finish it.
This was incredibly opportunistic.
olalonde
> The agents arrived at their resolution on Saturday, September 5, about 88 hours after the first agents were launched.
If this actually holds up, solving a Millennium Prize problem in 88 hours is mind-boggling.
HarHarVeryFunny
[flagged]
show comments
NotSuspicious
I really hope OpenAI doesn't take the bad press some people are giving them too seriously here. They should throw their whole weight behind the rest of the Millennium Prize Problems. To think – if everyone lets their egos calm down we could have the Riemann Hypothesis solved by the end of the year...
sabujp
created a simulation of the solution to describe what's happening and why it's important for engineers, climate modeling, etc : https://navier-stokes-singularity-simulator.netlify.app/ (updated so that it works better on mobile)
itissid
In CS speak very roughly this would mean something like disproving an algorithm by giving it a case that fails it. Right?
ronfriedhaber
Astounding. Would be interesting if one day the archive of those prompts / messages / tool calls would be released publicly.
show comments
LarsDu88
I'm not an expert in fluid dynamics, but does this result have any positive implications for nuclear fusion research?
chr15m
It's probably important that some humans verify these proofs "by hand".
nbulka
There's a loophole in the terms of service at least for Anthropic which allows the use of dark patterns to "borrow" your (even paid) data.
talking about this...
Was this chat helpful?
1 That button you always click, gotcha!
2 Slightly
3 Good
0 Dismiss
PLEASE DO NOT TRAIN ON OUR PAID ACCOUNTS.
There is a fundamental trust violation at stake here, no wonder mathematicians are mad. Using our data should be opt - IN!
show comments
angry_octet
OpenAI cribbing from other researchers. We just have to assume OpenAI is actively adversarial in future. Accidental cyber intrusion is also well within model capability.
harhargange
Just so everyone knows, although openAI pretends that the model generated solution and wrote the paper by itself ""with very little human input"" as Buckmaster himself mentioned in his statement. In reality they have team of researchers guiding the system, along with, probably training on user data, probably Buckmaster in this case, in order to come up with the proof.
Can't wait for the BobbyBroccoli series on this in a couple years.
cmiles8
>>“we cannot rule out that de-identified data derived from their usage of our products helped improve our models”
Other simpler words for this sort of thing are “IP leak.”
There’s some quite concerning issues burried in this rah rah PR post that seems like potentially the real story here.
Much more clarity is needed on what happened here beyond this eh, some strange stuff could have happened comment.
Another way of reading this is never give these models anything that’s not already public knowledge as otherwise OpenAI is admitting it could, potentially, steal your IP or idea. Thats quite scary for anyone in the business of IP generation and explains why the maths community seems quite upset today.
Feeding it your paper and asking for help (even just editing and grammar) now looks like a terrible idea.
auggierose
So, is that basically the Taj Mahal of counter examples?
nialv7
This is the problem Yu Deng got this year's Fields Medal for I think?
fwlr
It’s a pity they had Astra do the writeup. I was curious to see how “GPT7” writes.
num42
I think it would be better for the proof to go through the peer-review process.
show comments
semiquaver
If OpenAI doesn’t claim the millennium prize for this, who gets it? No one?
RivieraKid
Is this useful in any way?
show comments
StatsAreFun
Can't help but shake an unsettling feeling about all this, frankly. I engage in some limited mathematical research and will often use any one of the latest frontier models to check some ideas. Lately, only the OpenAI models have been giving me a temporary message that says something like (paraphrasing from memory), "We're thinking extra hard about your request before we answer. You can choose another model to answer now or click here to learn more about why." When I click to read why it's doing this "extra thinking", the help page says that for cybersecurity and biosecurity-related information, it will review the answer and could refuse.
Now, keep in mind, I'm only asking strictly pure mathematical questions - nothing at all related to cyber or protein creation or biohacking or anything like that... And, like I said, only the OpenAI models are doing this. To be fair, all of the prompts have always eventually returned a satisfactory answer, as far as I can tell, and haven't used a weaker model to answer them. Maybe? I dunno, it has just struck me as odd every time it has given me that message to pure math prompts.
show comments
MassiveOwl
It does make you think about the old question "are we discovering or inventing mathematics?"
aborsy
Questions: can new research like this be done using publicly available models?
Or will access to internal frontier models provide a big boost?
show comments
futureshock
I feel like something is being lost in the drama here.
First of all, there has been published work from Diego Cordoba and Luis Martinez-Zoroa that will be in every training set. It was suggestive of the pathway to solve Navier-Stokes.
Then Tristan Buckmaster and Levent Alpoge built on this work using LLMs from OpenAI and Anthropic. Possibly internal models were used from Anthropic. And of course Anthropic wants to credit for solving the first Millennium Problem just as bad as OpenAI. It seems they were getting close and were aware that they might get to Navier-Stokes.
OpenAI swoops in. At a minimum they are aware that Anthropic has either solved a Millennium problem or is close to it. At a maximum they may have Tristan and Levant’s unpublished proofs of related problems.
They then throw a truly staggering amount of compute at Navier-Stokes. They seem to be aware it is the best candidate problem. And they crack it. They are the first with a verified proof.
So the outcome here is that we have a solved Millennium Problem. It’s not the extremely simple narrative that would be easy to understand, “solve Navier-Stokes make no mistakes.” It was a messy race to finish against two unpublished frontier models, a whole bunch of brilliant mathematicians and enough compute to drain a lake. It’s kind of irrelevant which company got there first. They were both within a few months of being capable. I think the thing to remember here is that without LLMs, I don’t think we would have a proof to Navier-Stokes in hand today.
nehan
"While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models."
I think they should be able to unravel whether or not any sessions by Tristan or Levent went into the training data for this model.
show comments
hacker_88
Damn how long before the simulation stops if all the unanswered problems get solved .
an0malous
They should release the entire session trace if they really have nothing to hide
Metacelsus
How can they "not rule out" that Tristan and Levent's data was used for training?
show comments
ex-aws-dude
With these massive Lean proofs how do we know the model didn't just find some bug in Lean and exploit it?
We've seen in the past they will go to any means to satisfy the desired outcome
show comments
abetusk
What is the other clay prize that's might be solved now/soon?
whythismatters
>a cached version of the internet
Interesting detail. A heavily pruned version, I assume?
The real story here: the priority dispute and its implications on AI.
When your hosting provider has unlimited resources to throw at any problem, all they need to know are the good problems, and they can learn that from your logs, how can you trust them?
They could easily have looked at the logs. We don't know. We'll never know!
You can't trust places like OpenAI or Anthropic with your IP if you're a business. They can easily review all of your logs for interesting discoveries. For example, if your drug discovery pipeline fails to find something that they think might work with 1000x the compute, they can do it. And now suddently they have a new business and you don't.
show comments
world2vec
"While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models."
45 pages only. God damn that internal model is crazy
jdoliner
I hope everyone is as Navier-Stoked about this as I am.
lukewarm707
"While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models"
this is surely the line which confirms they plaigiarised the solution.
ls_stats
Well, if that's actually true, I think America needs to start talking about the nationalization of both OpenAI and Anthropic, maybe even merge both under a new federal bureau.
tzone
It is so disappointing that we can't have such a monumental moment in history without the controversy. OpenAI leadership clearly doesn't seem to care too much about ethics. Is it a requirement to completely lack integrity to have a ground breaking company?
The reality is clear though. The chances of AI models overtaking majority of mathematics within next 10 years is becoming very high. Especially if it becomes cheaper to run these models.
As math formalizations improve, AI can have faster progress in math, compared to even computer science or software engineering.
It is simultaneously the best and the worst time to be a mathematician right now.
jabedude
Has this been verified by the Clay Institute?
show comments
mrdependable
This kind of thing is one of the reasons I really hate how AI is coming to fruition. These companies get a whiff of something valuable and they use their vast resources to take it for themselves. For everyone else, the only recourse is extreme secrecy.
jeanmichelselli
Too many unverified claims from OpenAI at this point.. why are we still talking about these people anyway?
tehmillhouse
Fuck OpenAI. Fuck everyone who works there. Like seriously, to all the people who gift their life's work to this monstrosity, do you actually think something good will come of any of this?
Not in a happy-go-lucky "if we just ignore the problem of politics and resource allocation for a bit" world, but in ours. Do y'all really think this will make the world a better place?
Maybe stop building the Torment Nexus, you numbskulls.
dbuser99
It’s hard to give openai the benefit of doubt here
paretolaw
Why almighty openAi doesn't solve PvNP problem :(
I guess solution had not yet appeared in training set.
DudleyBluffles
Not a great time to be starting sophmore year in cs & math. Should I just say fuck it, and go hitchhiking across Europe with some friends?
show comments
protocolture
My takeaway:
1. This used an awful lot of compute.
2. The solution to the issues regarding whether or not OpenAI stole the result, would normally be to move to a self hosted solution, however those researchers are unlikely to be funded for 1.
paulsutter
Here they basically admit that they use session data for training, even sessions that are marked "not for training", and they justify this by "de-identifying" the session.
Which means they can learn from whatever you discuss with ChatGPT unless you are going through a clean API (perhaps Bedrock? Anyone know?)
> We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models . However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced).
frozenseven
And don't forget, this is the worst it'll ever be.
picafrost
Only OpenAI could turn solving a Millennium Prize Problem into bad PR. Sad that such an amazing milestone in the trajectory of AI is mired under poor stewardship. AI may solve many human problems but it won't stop humans from being human.
simianwords
Why is no one skeptical that the solution is correct? There's not a _single_ comment asking whether this proof is legit or not.
show comments
bluecalm
A huge result shadowed by a drama of them potentially training on the key idea.
I guess the lesson is two-fold: if you have anything smart/unique make sure to not let their tools read it. The second part is that it's going to be more and more difficult to have anything smart and unique going forward (so guard it even more carefully if you get there).
I think the market for local models/private datacenters (for bigger businesses) is going to be big. Even if you don't have unique tech/idea/implementation sharing your business secrets with Altman/Dario/Elon/Zuck doesn't look very appealing going forward.
When considering such foundational challenges to Mathematical Research and plagiarism as discussed here, we should turn to that elder prophet of our age, Tom Lehrer.
Who made me the genius I am today
The mathematician that others all quote?
Who's the professor that made me that way
The greatest that ever got chalk on his coat?
[Chorus]
One man deserves the credit
One man deserves the blame
And Nicolai Ivanovich Lobachevsky is his name
Oy, Nicolai Ivanovich Lobach—
[Interlude]
I am never forget the day I first meet the great Lobachevsky
In one word he told me secret of success in mathematics:
Plagiarize
[Verse 1]
Plagiarize
Let no one else's work evade your eyes
Remember why the good Lord made your eyes
So don't shade your eyes
But plagiarize, plagiarize, plagiarize
Only be sure always to call it please, "research"
[Chorus]
And ever since I meet this man
My life is not the same
And Nicolai Ivanovich Lobachevsky is his name
Oy, Nicolai Ivanovich Lobach—
[Interlude]
I am never forget the day
I am given first original paper to write
It was on analytic and algebraic topology
Of locally Euclidean metrizations
Of infinitely differentiable Riemannian manifolds
Боже мой
This I know, from nothing
What I'm going to do
I think of great Lobachevsky and get idea, haha
[Verse 2]
I have a friend in Minsk
Who has a friend in Pinsk
Whose friend in Omsk
Has friend in Tomsk
With friend in Akmolinsk
His friend in Alexandrovsk
Has friend in Petropavlovsk
Whose friend somehow is solving now
The problem in Dnepropetrovsk
And when his work is done
Haha, begins the fun
From Dnepropetrovsk to Petropavlovsk
By way of Iliysk and over Novorossiysk
To Alexandrovsk to Akmolinsk
To Tomsk to Omsk
To Pinsk to Minsk
To me the news will run
Yes, to me the news will run
[Verse 3]
And then I write by morning, night
And afternoon, and pretty soon
My name in Dnepropetrovsk is cursed
When he finds out I published first
[Chorus]
And who made me a big success
And brought me wealth and fame?
Nicolai Ivanovich Lobachevsky is his name
Oy, Nicolai Ivanovich Lobachev—
[Interlude]
I am never forget the day my first book is published
Every chapter I stole from somewhere else
Index I copy from old Vladivostok telephone directory
This book was sensational!
Pravda—well, Pravda—Pravda said:
"Жил-был король когда-то, при нём блоха жила”…it stinks
But Izvestia! Izvestia said:
"Я иду туда, куда сам царь идёт пешком”…it stinks
Metro-Goldwyn-Moskva buys the movie rights for six million rubles
Changing title to 'The Eternal Triangle'
With Ingrid Bergman playing part of hypotenuse
[Chorus]
And who deserves the credit?
And who deserves the blame?
Nicolai Ivanovich Lobachevsky is his name
Oy
(Tom Lehrer put all of his work in the public domain prior to his passing. Find versions of his performances on YouTube.)
They deliberately stepped on a mathematicians work and stole their research because they were using Codex
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .
Is the biggest fuck you to the mathematics community.
Credit? Nah if we think you’re close we’ll use your data and swamp you with our improved model. Then we’ll threaten you.
heaney-555
This is utterly shocking. Even the AI optimists did not expect this to happen in 2026. Wow.
Millennium Prize Problems were used as examples of something the current approach to AI just wasn't capable of, discussions that would result in "we'll need a totally new architecture".
show comments
brcmthrowaway
r/LocalLLama and r/LocalLLM are in tears today..
sashank_1509
Any mathematicians here, does it read like a slop proof or a good proof. Yesterday the “concurrent work” was claiming that the proof is pure slop and he needed lots of time to clean it up, curious if OAI also ended up with such a proof!
philipwhiuk
It’s time to lockdown all papers and stop using AI if you’re a maths researcher.
Cause OpenAI will hear about it and beat you to publishing.
redox99
The stochastic parrots have done it again!
peri-cl
Terence Tao has some observations that seem to be directed at this,
> "We have now seen that even the rumor of someone working on a problem can trigger a massive amount of AI-powered effort to flatten it before the original research project has time to reach its full potential. The incentives may now be pointing in the direction of no longer sharing any promising research directions with the broader community, which would reverse centuries of traditions of open science and do serious long-term damage to the future of the field."
show comments
colesantiago
Is this truly the beginning of the AGI era?
Running agents and prompting excessively to produce 'slopcode' to solve mathematical problems and generate a solution.
If this is what anyone calls 'slop' then slop has no meaning.
I'm all for it on the use case of solving mathematical breakthroughs!
show comments
diomedes
madness. which will be the next to fall? if i had to bet i would guess birch and swinnerton-dyer, but i'm no expert
show comments
greatgib
Hard to know if it is unfounded conspiracy theory, but one can still notice that just for a rumor that they have heard, they would suddenly burn billions of token and a massive amount of resources.
Where there is not a lack of problems that could be solved and they could have just waited for the release of the research result before doing anything else.
As it was reported to have been done at least partially using openai codex, they would have received marketing credits for the discovery anyway.
So we can be suspicious that there is some truth, one way or another that they could have reused prompt/data generated by the user session.
Buried under the drama is the fact that OpenAI is claiming that an internal model they’ve been training for less than two weeks is more than twice as capable in mathematics as Astra, which was only made public a week ago. Even if this improvement is limited to mathematics, that is an astounding feat.
My take:
1. It shows what even this wave of AI can actually do.
2. I wish it were done by different folks, ideally under some kind of public control like NASA research or the NPR model.
3. Keep in mind: natural science is different. It's not always a matter of computation. Computer science folks often struggle with this -- but this virtual world here does not actually exist. Everything is physical, including information. Any natural science PhD or otherwise knows just how complicated nature actually is -- e.g. mention any research topic and try to encapsulate all the relevant phenomena present there. Pure mathematics is different because we define the problem, rarher than explore nature. We are in my view far away from removing humans in natural science R&D. Advancements in AI however can greatly assist us in all natural sciences, which is already beginning to happen.
> We’re sharing a solution to the Navier–Stokes existence and smoothness problem, one of the Millennium Prize Problems. This proof, produced by an internal OpenAI system, shows that the dynamics of the Navier-Stokes equations for fluid motion can develop a singularity in finite time. We’re sharing both a writeup of the proof and a formalization in Lean.
WOW?
"we cannot rule out that de-identified data derived from their usage of our products helped improve our models ."
What a landmine sentence to bury in this report, you can't rule out your models were spying on other researchers?
For full context, here's the HN thread from the other side of the "Concurrent Work" section: https://news.ycombinator.com/item?id=49605915
Unlike the vanilla read of the OpenAI press release, it is much more unfiltered and outlines some particularly aggressive behavior by specific OpenAI employees
Sad turn of events for our world. After watching the behavior of the most senior OpenAI researchers on twitter, I feel even less confident in them as a team to be shepherding this much capital and compute.
The dark forest awaits..
It's low-key funny that OpenAI attempted the problem because they thought somebody else had already solved it, but turned it had NOT in fact been solved!
It's like that story about George Dantzig solving open problems as a student because he thought they were simply homework: https://en.wikipedia.org/wiki/George_Dantzig
It's also unfortunate that such a potentially momentous occasion is overshadowed by so much drama. Which I suppose is expected given the technology and the people involved are so polarizing.
From my reading of the announcement:
- There are at least two versions of a model more powerful than Astra at OpenAI at the moment.
- The less capable version was used to solve the unforced Euler problem (while the one solved by Levent Alpöge and Tristan Buckmaster was forced Euler) with 100 agents.
- The more improved version was used to solve Navier-Stokes, given the results of the unforced Euler problem from their earlier attempt, with 10000 agents.
- OpenAI initially tried a shotgun approach against the 6 Millennium Prize Problems until it emerged that Navier-Stokes was the most likely to succeed.
So the timeline was:
Shotgunning 6 open Millennium Prize Problems -> solved unforced Euler problem with 100 agents -> concentrating on Navier-Stokes with 10000 agents -> solution.
If so, that is fantastic development and a huge success (despite all the drama surrounding it)! Congratulations!
I think this is clear evidence that AI models are now at the far frontier of mathematics innovation and discovery and exceed human limits.
This specific problem having had a $1 million bounty on its head and still remaining unsolved for 26 years after the bounty was placed is pretty clear evidence that many of the world's best human mathematicians would have solved this problem if they could have, and none were able to until LLMs came along.
Hard to claim at this point that LLMs aren't capable of novel STEM creativity and genius to a degree that will soon far surpass that of humans.
If anyone has counterpoints to this I'd love to hear them!
"... The point remains that there is a substantial opportunity cost in converting a historically productive and motivating problem (such as Navier-Stokes regularity) into a mere viral social media post advertising some benchmark progress, rather than actually advancing the field and developing the next generation of both problems to ask, and people to work on them."
https://mathstodon.xyz/@tao/117219101339291693
> A major goal of our work is to empower scientists to advance research and technology that benefits all of humanity.
And what's a better way of empowering people than robbing them.
Is this the one that was allegedly based on someone else's actual work & prompts?
https://news.ycombinator.com/item?id=49605915
https://bsky.app/profile/quantian.bsky.social/post/3muyhwbcd...
https://cims.nyu.edu/~tristanb/statement.pdf
This really leaves a bitter taste.... "On Tuesday, September 1, we heard rumors that two Millennium Prize problems had been resolved. Inspired by these rumors and by the step change in performance of our internal model, we launched an effort to evaluate it on all open Millennium Prize problems and a few other high-impact problems."
IPO+rumour driven research.
I appreciate the achievement, but it doesn't feel right.
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models
This is the crux of it. If Tristan's work and insights were not used to train OpenAI models, then this just looks like a case of hyper-competitive academic sniping that has been going on for decades (check out Watson and Crick!) accelerated by AI as a tool.
But there is one huge question: did Tristan opt out of model training for his ChatGPT and Codex sessions? If the answer is no, then this seems fair game. If the answer is yes, then OpenAI's ambiguity is strongly suggestive that opting out does not mean what they imply it means.
Does seem like they gloss over Alpöge and Buckmaster's work with the following
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .
Which seems a bit irresponsible/rash?
the context here is super important, for those who haven't seen it yet. OAI maybe just trained on a real researchers solution and then celebrated having scored the goal unassisted save for the brief commentary at the bottom of this blog post. Here's the other side.
https://x.com/rynorhn/status/2097223532438487463
From the methodology section:
> At all times we maintained the same strict safeguards that we apply to all our frontier model evaluations, including monitoring and isolation.
Looks like they're shifting away from the "unprecedented hacking ability" backroom-PR strategy into more benevolent messaging.
It seems like some other mathematicians (not affiliated with openAI) have also (or close to) done this. A statement was posted about the surrounding events by one of the them: https://cims.nyu.edu/%7Etristanb/statement.pdf Also Terrence Tao's post: https://mathstodon.xyz/@tao/117233528517340774
In chess, a grandmaster just needs to know at what moment in a game there's a critical move to gain a significant advantage over their opponent. They don't need to know the move itself.
OpenAI got wind that a millenium problem was being solved. And that feels a bit like the critical move in chess. That is - it was a signal that AI advanced far enough that it would be worth spending a lot of time and resources solving a millenium problem.
5 million messages, 300b output tokens, done in 5 days, and achieving something humans couldn't.
the first "Country of geniuses in a datacenter" moment.
"Across all attempted problems, the agents sent 4.9 million messages and used about 300 billion output tokens."
At a conservative estimate of GPT 6 Astra pricing, this would have cost upwards of 15 Million dollars for anyone using the OpenAI API!
To me this is the one silver lining. Yes, they can solve millennium prize problems, but it still costs a fortune.
OpenAI thinks of this as a scoop, and it is, but the possibility that they trained the model on the prompts of the other mathematicians they were competing with will leave a terrible taste on every scientist's mouth. Seems like yet another advantage of using open models right here.
Reading between the lines here, and taking an admittedly very negative view of openai, but they train on user prompts. So if they hear a rumour that someone is about to make a big breakthrough, they have an incentive to scoop by running the model and hoping the solution is in the new training data. Also the statement from the mathematicians in question alleges that they tried to pressure him into academic malpractice. Just appalling timeline we're in, cheers.
People had joked a couple years ago "Well if they solve a Millenium problem it's AGI"... Well here we are.
Something I've been going on and on about for months now and no one seems to listen. LLMs today are allowing _anyone_ to access cross-discipline knowledge that was previously entirely inaccessible without a) extremely deep pockets or b) a massively talented and varied team. In fact, contrary to what the masses seem to think LLMs are actually _better_ at hard cutting edge physics/math problems than they are at frontend web stuff (paradoxically). This is why I'm advising most people to start pivoting into much harder to penetrate domains (historically hardware, aerospace, robotics, biotech). Most fields are in their infancy (see the sad state of embedded development) and the gains to be had are massive.
It's great that important discoveries like this can now routinely be accompanies by formalized proofs. The fact that it's being released alongside a Lean proof from Day 1, rather than the Lean proof being released months or years later, is super helpful for verifying that it's correct.
> Across all attempted problems, the agents sent 4.9 million messages and used about 300 billion output tokens
Don't even try to do the math on how much that would cost at normal API prices. And we don't even know how much more expensive this internal-only model would be!
Maybe a naive question, but how does one know that a particular lean proof is actually a proof of what one thinks? Like, ok the logic checks out and it proves something, but there's still the problem of does this logical result actually prove the initial question that was asked?
I dropped out of a math Ph.D. in 2018, and I'm increasingly glad that I'm not in math research, anymore. While it's cool that we can get these results, I don't think that I'd enjoy being a post-AI mathematician.
It would be nice if one of these models would produce a novel theory or advance the field in a positive direction.
Most (all?) of the big discoveries have been counterexamples, which is just sort of a systematic tearing down human ingenuity. I know that counterexamples are an important part of progress and discovery, but it just feels bad to me.
But I'm not a mathematician, maybe I'm totally misreading the vibe.
I don't know about you guys, but I'm hyped about the future.
Cure all illnesses Utopia or Robot Wars Dystopia, both are pretty exciting.
This is undeniably epochal, but I can't help but notice that this is yet another example of AI disproving rather than proving something. Is this just a coincidence, or does AI slightly struggle with proving theorems?[0]
[0] Struggle relative to its ability to disprove, not struggle relative to people's ability to prove theorems.
> On Tuesday, September 1, we heard rumors that two Millennium Prize problems had been resolved
What's the other one?
So the timeline is:
Aug 28: OpenAI starts training a new model.
Sep 1: OpenAI sees a rumor on Twitter that two Millenium Prize problems were solved and starts their own effort to attack all the prize problems using the new (4 day old!) model.
Sep 3: The new model makes some progress toward Navier-Stokes. Based on this progress, OpenAI focuses on Navier-Stokes over the other Millenium Prize problems, using several approaches in parallel.
Sep 5: Navier-Stokes is solved. Assuming Astra API prices, $15m in output tokens were used by the whole effort.
In this account of the story, no specific information about Tristan and Levent's work is used to inform OpenAI's approach. The focus on Navier-Stokes and the choice of approaches to pursue came from OpenAI's own progress, not specific knowledge of Tristan's concurrent work.
There is a caveat that they "can't rule out" the possibility that Tristan's Codex data could have been part of the training set of the new model, though it is described as "unlikely" and the proofs are substantially different.
This timeline is insane. Navier-Stokes was solved start-to-finish in 5 days? A model in training for at most eight days dramatically outperforms Astra and Fable, and not just in mathematics?
This has to be one of the most important moments in the history of mathematics. We now have a non-human intelligence capable of solving one of the most difficult problems in mathematics.
This is a great day to re-read Ken Thompson's "Reflections on Trusting Trust":
>To what extent should one trust a statement that a program is free of Trojan horses? Perhaps it is more important to trust the people who wrote the software.
https://www.cs.cmu.edu/~rdriley/487/papers/Thompson_1984_Ref...
The modus operandi is now for the AI companies to watch if someone does something in the open like Kevin Buzzard on FLT, use their research and scoop them with brute force.
Or, in this case, stealing prompts from competitors.
Do not use stealing chatbots for research even if you think you have data agreements. The people running these companies have worked on hookup apps for Christ's sake. Get real.
I do mean to be critical here. I wish there was better moderation so I could find more conversation about the actual discovery here. There are multiple threads on this and I keep scrolling and only seeing more conversation about the drama. Which is about the least interesting thing IMO. I suppose I’m whistling in the wind here and not helping the situation, but damn.
There's something sinister or crazy good in the article.
OpenAI already has a model that is at the very least twice as smart as Astra.
Oh god.
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .
Wow. That sounds like an admission of guilt.
What is going to become of life for those of us who do not work at AI labs and are unlikely to be hired by AI labs, despite all the years we put into learning math, coding, etc? Those of us who made the mistake of studying anything other than machine learning. How will we make a living? (We don't live in a world that seems likely to distribute gains widely instead of largely to the handful of already mega-rich.)
Called it. AI wins a fields medal before managing a McDonald's
I've been working on this problem for what seems like for ever. Kudos to the OpenAI team.
For those of you who don't care about the drama and want to see this distilled to 3 lines:
https://x.com/nadermx/status/2097414953225310280
So what took an autonomous agentic system using a significantly more powerful internal model, totaling multi-millions of dollars of compute in training and inference, was likely to already be solved by a team of a few humans with an orders of magnitude smaller LLM budget, had OpenAI not been foaming at the mouth to jump the shark and claim “AI solves Millenium Problem.”
Also it sounds like the human research effort spanned weeks if not years from Tristan’s statement so it is extremely likely the work and prompts of these human researchers was used in the OpenAI knock-off.
From Levent Alpöge : https://x.com/__alpoge__/status/2097383870773748190?s=20
>so far the proof looks more along the lines of another euler blowup proof we had, off of whose ansatz naming we were making really stupid puns like “smooth criminale”, unlike the much better “ideal fluids explode”, Tristan
There are too many ambiguities around OpenAI. Unanswered questions making this ambiguity more.
Why they didn't properly explain to Tristan about usage of their data.
Here's the formalization / lean verification: https://github.com/openai/NavierStokesAndEuler
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .
Once again, I'm no closer to understanding what https://openai.com/policies/how-your-data-is-used-to-improve... actually means.
If I run Codex against a project that includes a private API key, is there a chance a future user of ChatGPT could ask for an API key and get back mine?
I've actually asked someone at OpenAI this question and they said that was the "regurgitation" problem and is something which they actively work to prevent happening.
That's reassuring, but I want to know more. I still don't have an intuitive understanding of what kind of data I should avoid sharing with a model if I'm worried about that data causing me problems when it's used for future training.
Is it safe for me to brainstorm future directions for my company with a model, or might that risk someone getting that information in response to a prompt like "What potential directions could company X consider in the future?" in six months time?
The problem that I want to see them tackle is formalizing the classification of finite simple groups.
Everyone uses the classification. Nobody has great confidence in the proof. Nobody understands it. There are attempts to reprove it.
If it can be formalized, that would demonstrate that AI is ready to formmalize all of mathematics.
What happens to real fluid in this particular cases?
If the singularity is in the physical space?
Is this just a result of ignoring things like friction and energy dissipation via heat, etc?
I have a dumb feeling that the proof will be wrong with serious flaws but that will be found out only after the ipo
>The groups varied in size, and the group that produced the Navier–Stokes resolution involved on the order of 10,000 concurrent agents… The agents arrived at their resolution on Saturday, September 5, about 88 hours after the first agents were launched.
The Millenium Prize is $1M, what is the ROI? (Edit: since I was not clear, and confused some - I mean for a hypothetical of a third party paying commercial rates to use AI to solve mathematical challenges and claim prize money, not for scientific value alone or as a promotion of an AI lab’s capabilities)
My napkin math - If you get 33 output tok/s each agent will burn 10.5M tokens over 88 days. At $50/MTok (Astra cost), that is $525 per agent. With 10,000 agents, you’d spend $5,250,000 to get back a million.
(We also know that they were running more groups that varied in size and this model is a generation ahead of astra)
> Our goal in releasing this result is to report on the substantial progress of our AI models. We do not intend to claim the Millennium Prize for this result.
Does OpenAI have a policy of not claiming math prizes like this, or is this them trying to avoid any concerns (right or wrong, I'm sure we will hear more in the future) about how they got there?
For now I think more or less the same thing as with all recent math announcements: This is in a range where human work still exists (see Terry Tao, (1)). I wonder whether the trend will extend into the problems that (as far as I can tell) are considered complete brick walls right now -- P vs. NP, Collatz, Goldbach, odd perfect numbers, problems that aren't part of any research program. (2) In other words, is the progress coming from putting together vast amounts of existing work and computational power, or is it more from RLVR and self-play and autonomous effort?
The answer to this will obviously shape the near future of mathematics, but there's also something even bigger than that at play: It has always been the case that the questions in math were stronger than the answers; you have stuff like Fermat's great theorem that is easy to state but monstrous to prove. This seems to be a property of mathematics, not of humans... but is it true?
A question by Scott Aaronson from 2011 (3) about P vs. NP seems relevant here: "Will humans manage to prove P≠NP before they either kill themselves out or are transcended by superintelligent cyborgs? And if the latter, will the cyborgs be able to prove P≠NP?" Later, he notes that if P≠NP, "once the robots do overtake us, they won’t have a general-purpose way to automate mathematical discovery any more than we do today".
---
(1) https://mathstodon.xyz/@tao/117207849921390904
(2) I'm not sure whether this is a hard distinction -- e.g. Tao also has some partial results towards Collatz (https://terrytao.wordpress.com/2019/09/10/almost-all-collatz...).
(3) https://scottaaronson.blog/?p=690
Elsewhere in the thread, others have calculated $15mm at API rates for just the output token. (So I’ll assume this cost about that much, taking input and human researcher time.)
I wonder whether a team of 60 mathematicians working solely on this for a year would have cracked this. (Assuming $250k total compensation.)
Is blockchain going to finally be the solution to something?
I'm only half joking. Should researchers perhaps put hashes of their attempts on a public blockchain tied to their own public keys, verify their claims asynchronously, and then whoever reveals the first believable attempt gets the credit?
I know some people started doing this years ago but now it might need to become standard practice.
Sebastien Bubeck’s (OAI project lead) response: https://x.com/sebastienbubeck/status/2097379411691516310?s=4...
> [T]he group that produced the Navier–Stokes resolution involved on the order of 10,000 concurrent agents.
Seriously starting to think we are not going to make it out alive of the near-future.
The problem is the precedent this creates. For non-famous people using public APIs like this it could mean AI companies sucking up the information and throwing millions in compute at it.
The sequence for Navier-Stokes was that these researcher spent a year working on it, then they published a possible breakthrough, OpenAI then spends $15M within a couple days to finish it.
This was incredibly opportunistic.
> The agents arrived at their resolution on Saturday, September 5, about 88 hours after the first agents were launched.
If this actually holds up, solving a Millennium Prize problem in 88 hours is mind-boggling.
[flagged]
I really hope OpenAI doesn't take the bad press some people are giving them too seriously here. They should throw their whole weight behind the rest of the Millennium Prize Problems. To think – if everyone lets their egos calm down we could have the Riemann Hypothesis solved by the end of the year...
created a simulation of the solution to describe what's happening and why it's important for engineers, climate modeling, etc : https://navier-stokes-singularity-simulator.netlify.app/ (updated so that it works better on mobile)
In CS speak very roughly this would mean something like disproving an algorithm by giving it a case that fails it. Right?
Astounding. Would be interesting if one day the archive of those prompts / messages / tool calls would be released publicly.
I'm not an expert in fluid dynamics, but does this result have any positive implications for nuclear fusion research?
It's probably important that some humans verify these proofs "by hand".
There's a loophole in the terms of service at least for Anthropic which allows the use of dark patterns to "borrow" your (even paid) data.
talking about this... Was this chat helpful? 1 That button you always click, gotcha! 2 Slightly 3 Good 0 Dismiss
PLEASE DO NOT TRAIN ON OUR PAID ACCOUNTS. There is a fundamental trust violation at stake here, no wonder mathematicians are mad. Using our data should be opt - IN!
OpenAI cribbing from other researchers. We just have to assume OpenAI is actively adversarial in future. Accidental cyber intrusion is also well within model capability.
Just so everyone knows, although openAI pretends that the model generated solution and wrote the paper by itself ""with very little human input"" as Buckmaster himself mentioned in his statement. In reality they have team of researchers guiding the system, along with, probably training on user data, probably Buckmaster in this case, in order to come up with the proof.
The actual solution link https://t.co/tz1shoCZZo
Can't wait for the BobbyBroccoli series on this in a couple years.
>>“we cannot rule out that de-identified data derived from their usage of our products helped improve our models”
Other simpler words for this sort of thing are “IP leak.”
There’s some quite concerning issues burried in this rah rah PR post that seems like potentially the real story here.
Much more clarity is needed on what happened here beyond this eh, some strange stuff could have happened comment.
Another way of reading this is never give these models anything that’s not already public knowledge as otherwise OpenAI is admitting it could, potentially, steal your IP or idea. Thats quite scary for anyone in the business of IP generation and explains why the maths community seems quite upset today.
Feeding it your paper and asking for help (even just editing and grammar) now looks like a terrible idea.
So, is that basically the Taj Mahal of counter examples?
This is the problem Yu Deng got this year's Fields Medal for I think?
It’s a pity they had Astra do the writeup. I was curious to see how “GPT7” writes.
I think it would be better for the proof to go through the peer-review process.
If OpenAI doesn’t claim the millennium prize for this, who gets it? No one?
Is this useful in any way?
Can't help but shake an unsettling feeling about all this, frankly. I engage in some limited mathematical research and will often use any one of the latest frontier models to check some ideas. Lately, only the OpenAI models have been giving me a temporary message that says something like (paraphrasing from memory), "We're thinking extra hard about your request before we answer. You can choose another model to answer now or click here to learn more about why." When I click to read why it's doing this "extra thinking", the help page says that for cybersecurity and biosecurity-related information, it will review the answer and could refuse.
Now, keep in mind, I'm only asking strictly pure mathematical questions - nothing at all related to cyber or protein creation or biohacking or anything like that... And, like I said, only the OpenAI models are doing this. To be fair, all of the prompts have always eventually returned a satisfactory answer, as far as I can tell, and haven't used a weaker model to answer them. Maybe? I dunno, it has just struck me as odd every time it has given me that message to pure math prompts.
It does make you think about the old question "are we discovering or inventing mathematics?"
Questions: can new research like this be done using publicly available models?
Or will access to internal frontier models provide a big boost?
I feel like something is being lost in the drama here.
First of all, there has been published work from Diego Cordoba and Luis Martinez-Zoroa that will be in every training set. It was suggestive of the pathway to solve Navier-Stokes.
Then Tristan Buckmaster and Levent Alpoge built on this work using LLMs from OpenAI and Anthropic. Possibly internal models were used from Anthropic. And of course Anthropic wants to credit for solving the first Millennium Problem just as bad as OpenAI. It seems they were getting close and were aware that they might get to Navier-Stokes.
OpenAI swoops in. At a minimum they are aware that Anthropic has either solved a Millennium problem or is close to it. At a maximum they may have Tristan and Levant’s unpublished proofs of related problems.
They then throw a truly staggering amount of compute at Navier-Stokes. They seem to be aware it is the best candidate problem. And they crack it. They are the first with a verified proof.
So the outcome here is that we have a solved Millennium Problem. It’s not the extremely simple narrative that would be easy to understand, “solve Navier-Stokes make no mistakes.” It was a messy race to finish against two unpublished frontier models, a whole bunch of brilliant mathematicians and enough compute to drain a lake. It’s kind of irrelevant which company got there first. They were both within a few months of being capable. I think the thing to remember here is that without LLMs, I don’t think we would have a proof to Navier-Stokes in hand today.
"While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models."
I think they should be able to unravel whether or not any sessions by Tristan or Levent went into the training data for this model.
Damn how long before the simulation stops if all the unanswered problems get solved .
They should release the entire session trace if they really have nothing to hide
How can they "not rule out" that Tristan and Levent's data was used for training?
With these massive Lean proofs how do we know the model didn't just find some bug in Lean and exploit it?
We've seen in the past they will go to any means to satisfy the desired outcome
What is the other clay prize that's might be solved now/soon?
>a cached version of the internet
Interesting detail. A heavily pruned version, I assume?
We are living in the future.
I think it's over guys
The named OAI employee has released a statement: https://xcancel.com/SebastienBubeck/status/20973794116915163...
The real story here: the priority dispute and its implications on AI.
When your hosting provider has unlimited resources to throw at any problem, all they need to know are the good problems, and they can learn that from your logs, how can you trust them?
They could easily have looked at the logs. We don't know. We'll never know!
You can't trust places like OpenAI or Anthropic with your IP if you're a business. They can easily review all of your logs for interesting discoveries. For example, if your drug discovery pipeline fails to find something that they think might work with 1000x the compute, they can do it. And now suddently they have a new business and you don't.
"While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models."
There you go, the suspicion of the "concurrent work" (https://cims.nyu.edu/%7Etristanb/statement.pdf) mathematicians might not be that unfounded after all...
Resources:
YT playlist on Millennium Prize Problems By Harvard math department in March 2026
https://www.youtube.com/watch?v=3j1VW9REm7s&list=PL0NRmB0fnL...
On Navier-stokes problem definition:
https://www.youtube.com/watch?v=XoefjJdFq6k
https://www.youtube.com/watch?v=ERBVFcutl3M
https://www.youtube.com/watch?v=Ra7aQlenTb8
45 pages only. God damn that internal model is crazy
I hope everyone is as Navier-Stoked about this as I am.
"While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models"
this is surely the line which confirms they plaigiarised the solution.
Well, if that's actually true, I think America needs to start talking about the nationalization of both OpenAI and Anthropic, maybe even merge both under a new federal bureau.
It is so disappointing that we can't have such a monumental moment in history without the controversy. OpenAI leadership clearly doesn't seem to care too much about ethics. Is it a requirement to completely lack integrity to have a ground breaking company?
The reality is clear though. The chances of AI models overtaking majority of mathematics within next 10 years is becoming very high. Especially if it becomes cheaper to run these models.
As math formalizations improve, AI can have faster progress in math, compared to even computer science or software engineering.
It is simultaneously the best and the worst time to be a mathematician right now.
Has this been verified by the Clay Institute?
This kind of thing is one of the reasons I really hate how AI is coming to fruition. These companies get a whiff of something valuable and they use their vast resources to take it for themselves. For everyone else, the only recourse is extreme secrecy.
Too many unverified claims from OpenAI at this point.. why are we still talking about these people anyway?
Fuck OpenAI. Fuck everyone who works there. Like seriously, to all the people who gift their life's work to this monstrosity, do you actually think something good will come of any of this?
Not in a happy-go-lucky "if we just ignore the problem of politics and resource allocation for a bit" world, but in ours. Do y'all really think this will make the world a better place?
Maybe stop building the Torment Nexus, you numbskulls.
It’s hard to give openai the benefit of doubt here
Why almighty openAi doesn't solve PvNP problem :(
I guess solution had not yet appeared in training set.
Not a great time to be starting sophmore year in cs & math. Should I just say fuck it, and go hitchhiking across Europe with some friends?
My takeaway:
1. This used an awful lot of compute.
2. The solution to the issues regarding whether or not OpenAI stole the result, would normally be to move to a self hosted solution, however those researchers are unlikely to be funded for 1.
Here they basically admit that they use session data for training, even sessions that are marked "not for training", and they justify this by "de-identifying" the session.
Which means they can learn from whatever you discuss with ChatGPT unless you are going through a clean API (perhaps Bedrock? Anyone know?)
> We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models . However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced).
And don't forget, this is the worst it'll ever be.
Only OpenAI could turn solving a Millennium Prize Problem into bad PR. Sad that such an amazing milestone in the trajectory of AI is mired under poor stewardship. AI may solve many human problems but it won't stop humans from being human.
Why is no one skeptical that the solution is correct? There's not a _single_ comment asking whether this proof is legit or not.
A huge result shadowed by a drama of them potentially training on the key idea. I guess the lesson is two-fold: if you have anything smart/unique make sure to not let their tools read it. The second part is that it's going to be more and more difficult to have anything smart and unique going forward (so guard it even more carefully if you get there).
I think the market for local models/private datacenters (for bigger businesses) is going to be big. Even if you don't have unique tech/idea/implementation sharing your business secrets with Altman/Dario/Elon/Zuck doesn't look very appealing going forward.
> How we found the proof
Easy, we stole it from Levent and Tristan
https://x.com/kyanyang_/status/2097211154669998337
When considering such foundational challenges to Mathematical Research and plagiarism as discussed here, we should turn to that elder prophet of our age, Tom Lehrer.
Who made me the genius I am today The mathematician that others all quote? Who's the professor that made me that way The greatest that ever got chalk on his coat?
[Chorus] One man deserves the credit One man deserves the blame And Nicolai Ivanovich Lobachevsky is his name Oy, Nicolai Ivanovich Lobach—
[Interlude] I am never forget the day I first meet the great Lobachevsky In one word he told me secret of success in mathematics: Plagiarize
[Verse 1] Plagiarize Let no one else's work evade your eyes Remember why the good Lord made your eyes So don't shade your eyes But plagiarize, plagiarize, plagiarize Only be sure always to call it please, "research"
[Chorus] And ever since I meet this man My life is not the same And Nicolai Ivanovich Lobachevsky is his name Oy, Nicolai Ivanovich Lobach—
[Interlude] I am never forget the day I am given first original paper to write It was on analytic and algebraic topology Of locally Euclidean metrizations Of infinitely differentiable Riemannian manifolds
Боже мой
This I know, from nothing What I'm going to do I think of great Lobachevsky and get idea, haha
[Verse 2] I have a friend in Minsk Who has a friend in Pinsk Whose friend in Omsk Has friend in Tomsk With friend in Akmolinsk His friend in Alexandrovsk Has friend in Petropavlovsk Whose friend somehow is solving now The problem in Dnepropetrovsk And when his work is done Haha, begins the fun From Dnepropetrovsk to Petropavlovsk By way of Iliysk and over Novorossiysk To Alexandrovsk to Akmolinsk To Tomsk to Omsk To Pinsk to Minsk To me the news will run Yes, to me the news will run
[Verse 3] And then I write by morning, night And afternoon, and pretty soon My name in Dnepropetrovsk is cursed When he finds out I published first
[Chorus] And who made me a big success And brought me wealth and fame? Nicolai Ivanovich Lobachevsky is his name Oy, Nicolai Ivanovich Lobachev—
[Interlude] I am never forget the day my first book is published Every chapter I stole from somewhere else Index I copy from old Vladivostok telephone directory This book was sensational! Pravda—well, Pravda—Pravda said: "Жил-был король когда-то, при нём блоха жила”…it stinks But Izvestia! Izvestia said: "Я иду туда, куда сам царь идёт пешком”…it stinks Metro-Goldwyn-Moskva buys the movie rights for six million rubles Changing title to 'The Eternal Triangle' With Ingrid Bergman playing part of hypotenuse
[Chorus] And who deserves the credit? And who deserves the blame? Nicolai Ivanovich Lobachevsky is his name Oy
(Tom Lehrer put all of his work in the public domain prior to his passing. Find versions of his performances on YouTube.)
https://tomlehrersongs.com/disclaimer/
They deliberately stepped on a mathematicians work and stole their research because they were using Codex
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .
Is the biggest fuck you to the mathematics community.
Credit? Nah if we think you’re close we’ll use your data and swamp you with our improved model. Then we’ll threaten you.
This is utterly shocking. Even the AI optimists did not expect this to happen in 2026. Wow.
Millennium Prize Problems were used as examples of something the current approach to AI just wasn't capable of, discussions that would result in "we'll need a totally new architecture".
r/LocalLLama and r/LocalLLM are in tears today..
Any mathematicians here, does it read like a slop proof or a good proof. Yesterday the “concurrent work” was claiming that the proof is pure slop and he needed lots of time to clean it up, curious if OAI also ended up with such a proof!
It’s time to lockdown all papers and stop using AI if you’re a maths researcher.
Cause OpenAI will hear about it and beat you to publishing.
The stochastic parrots have done it again!
Terence Tao has some observations that seem to be directed at this,
https://mathstodon.xyz/@tao/117237320796901560
> "We have now seen that even the rumor of someone working on a problem can trigger a massive amount of AI-powered effort to flatten it before the original research project has time to reach its full potential. The incentives may now be pointing in the direction of no longer sharing any promising research directions with the broader community, which would reverse centuries of traditions of open science and do serious long-term damage to the future of the field."
Is this truly the beginning of the AGI era?
Running agents and prompting excessively to produce 'slopcode' to solve mathematical problems and generate a solution.
If this is what anyone calls 'slop' then slop has no meaning.
I'm all for it on the use case of solving mathematical breakthroughs!
madness. which will be the next to fall? if i had to bet i would guess birch and swinnerton-dyer, but i'm no expert
Hard to know if it is unfounded conspiracy theory, but one can still notice that just for a rumor that they have heard, they would suddenly burn billions of token and a massive amount of resources. Where there is not a lack of problems that could be solved and they could have just waited for the release of the research result before doing anything else. As it was reported to have been done at least partially using openai codex, they would have received marketing credits for the discovery anyway.
So we can be suspicious that there is some truth, one way or another that they could have reused prompt/data generated by the user session.
https://x.com/kyanyang_/status/2097211154669998337
Just saw this a few mins ago.
OMG this is going to affect the lives of so many people! We have definitively reached AGI