I was a graph theory junkie long ago and even moved to Budapest for awhile to study among the greats. While I was there I started working on Barnette's Conjecture which came to occupy my thoughts over the next 24 years of my life, on and off as I worked in many different fields. Last summer I even thought for a few days that I had actually solved it.
But it's supposedly proven here - problem 180. I don't know what to think exactly. I spent thousands of hours on that problem. I really enjoyed it. Hearing that it is solved somehow makes me sad in a far-off way, like hearing an ex-girlfriend died suddenly in a car crash. I don't know, there's probably a lot of people feeling odd emotions tonight.
There's no Lean proof for this one so I'm digesting the paper. On the surface it looks like an approach I considered 24 years ago and abandoned.
I revisited the problem this summer, along with my partial solutions, when the previous round of stunning proofs came out. Several hours of work with Fable simply convinced me it wasn't yet solvable and reinforced how hard of a problem it was.
show comments
winfieldchen
> We prove the Unique Games Conjecture
The Unique Games Conjecture (sorry, "Unique Games Theorem" now!) is huge. It was a very significant pillar supporting many of the limits of the polynomial-time approximation algorithms in the graduate-level randomized and approximate algorithms course I took in theoretical computer science. Textbooks will have to be re-written.
With UGC proved, certain polynomial-time approximation algorithms used in difficult real-life problems are now known to be the best approximations we can achieve in polynomial-time:
> If UGC holds, the elementary algorithm that grabs both ends of an edge is fundamentally the best efficient algorithm that will ever exist. No amount of advanced linear programming or heuristics can achieve a ratio of 1.999.
> Under UGC, the Goemans-Williamson algorithm's 0.87856 ratio is mathematically optimal.
> UGC is considered the "Rosetta Stone" of approximation algorithms. In 2008, Prasad Raghavendra proved that for every single constraint satisfaction problem (CSP), a canonical Semidefinite Programming relaxation paired with the best rounding scheme achieves the optimal approximation ratio if and only if UGC is true. If the conjecture holds, the algorithmic boundary for an entire class of combinatorial problems is completely resolved.
Other hardness of approximation results from this UGC proof:
> [Max acyclic subgraph, a problem encountered in real life]: No polynomial-time algorithm can fundamentally outperform an unthinking coin toss.
> [Relative scheduling, another realistic problem]: As with acyclic subgraphs, the problem is "approximation-resistant": clever algorithms cannot beat random shuffling.
show comments
xanderlewis
As Kevin Buzzard recently said:
> In a 2020 piece in the Notices of the AMS, I asked the following question: “If one human had an understanding of all of modern pure mathematics simultaneously, how much further would they immediately be able to see?” Six years later we are beginning to understand the answer to this question.
show comments
qnleigh
It's notable that LLMs have now made substantial progress on four of the seven Millennium Prize problems, resolving one of them: Hodge, Birch-Swinnerton-Dyer, Riemann, and Navier-Stokes, which they resolved. No sign of P vs. NP or Yang-Mills existence and mass gap as far as I can tell, which is interesting.
Talking to some friends in physics this evening, most of the physics-related results that we could recognize were very mathematical, proving things rigorously where the physics community already had strong expectation. For instance, for a certain model of magnetism (the spin-1 Heisenberg chain), it was strongly expected that there is a finite energy gap between the ground state and the first excited state, but proving this rigorously was quite challenging. So while these are major results in mathematical physics, they probably don't rise to the level of a Millennium problem for the field.
It's interesting to think what a comparable breakthrough in physics might look like, since physics tends to favor things like conceptual understanding and applications over mathematical rigor. Maybe a new quantum algorithm, understanding of high-temperature superconductivity, a precise description of M theory...
show comments
rcr-anti
In Stellaris you can play as a civilization of robots who keep their biological creator race alive as "bio trophies". The bio trophies don't do anything meaningful besides by existing satisfy the need their ancestors placed in the robots to take care of them. Starting to wonder if that's the best we can hope for, if these things will be, if they aren't already, better than us at anything that matters.
show comments
enoether
Unique Games Conjecture [0] is a seminal conjecture in Complexity Theory, and is an underlying assumption for many, many inapproximability results. A valid proof is a big deal!
This includes a proof of Barnette's Conjecture, which is one of the graph theory conjectures that I tried attacking with SOTA models a few months ago. I like it because it is easy to understand with a basic knowledge of graph theory. I spent quite a bit of time on it and failed. Their proof looks approachable at first glance.
As a TCS/scheduling person, this one is definitely of lesser importance than UGC, but it has been an open problem since the book of Garey and Johnson in 1979:
A Polynomial-Time Algorithm for Three-Machine Unit-Job Scheduling [1]
Since some people talk about small numbers that pop up in integer multiplication results, here a completely different number appears:
Theorem 1.1. Let an explicitly listed finite directed acyclic graph specify the precedence constraints on n >= 1 nonpreemptive unit-length jobs on three identical machines. There is a uniform deterministic algorithm that constructs a feasible schedule of minimum makespan. Given also an integer deadline 1 <= T <= n, it decides feasibility exactly and returns a schedule whenever the answer is affirmative. Both tasks can be performed in O((L + 2)^150020) steps on a deterministic multitape Turing machine, where L is the total binary input length.
That is some crazy exponent -- plus an interestingly old computational model to boot; not something that is natural to most of us. I have no capacity to check its correctness today, but I hope it is true purely for the exponent.
An ex colleague of mine who is a world class mathematician recently got an ERC with ambitious goals to advance his field.
Literally every optimistic goal proposed to be worked on during this multi-year window has been solved in this one post. His and his entire group's work has just been done for him! They are all depressed as hell right now.
show comments
schleck8
Levent Alpöge (Anthropic mathematician) comment on the significance:
> Sure, mathematical history features a lot of incredible developments, like the invention of proof, zero, or the computer, and on the great problems our progress has been over timelines measured in decades or centuries. Obviously this technology didn’t appear today, but blurring our eyes a bit to combine the past ten years, with today a measurement of those developments, there is nothing comparable.
show comments
bcatanzaro
“I think at the heart of this issue is that humans have two competing natures: a tendency to compete and a capacity to appreciate beauty,” said Kai Shaikh, a graduate student in mathematics at the University of Toronto. “To me this seems to be a case of the former attempting to strangle the latter.” [1]
Beauty can be appreciated even when it is vast, even when it is beyond one's comprehension. I don't think this release should be primarily viewed as an outcome of competition. Instead it is revealing truths about the universe that were always there and always beautiful, even if we hadn't seen them yet. I believe there are infinitely more such beautiful truths currently hidden and waiting for us to discover.
This is significant progress and released without all the drama. Some very important progress in Reinmann, Hodge and unique games theorem. Point the repo to your agent and ask for the significance! In a way this is probably 50-100 years of math progress by humans
show comments
againstapples
As an AI "doomer" can I ask the non-doomer people here how you interpret the significance of results like these, and what kind of progress you expect to see in the next 1-5 years?
Like do you see the technology plateauing at the current level, do you expect progress will continue but only in mathematics, I'm interested to know why others are not concerned?
Beyond the reasoning capabilities of the unreleased model and that they can run a large number of agents in parallel, what makes me curious is that the writing of the proofs is quite human-readable. This is in contrast with the scientific text produced by the ChatGPT available to us, which writes horribly in a way that no human would write. One tell-tale is that they constantly attempt to be defensive and cover all edge cases like division by zero etc. that are clearly a non-issue for humans in some proofs or at least a human would not add this to the main statement, but AI is so overly careful that makes reading its proofs impossible. On the other hand, these new OpenAI proofs are very good.
foota
From their github: "The average result used the equivalent compute of roughly three hours of ChatGPT Pro thinking." That's pretty crazy.
show comments
kingstnap
Some of these are interesting ngl.
109. Integer multiplication below n log n
Surprising that this is possible.
158. The Euclidean plane cannot be colored with five colors.
Only 6 and 7 remain!
376. Universal computation in forced Navier–Stokes flows.
Let's hypothetically say I'm a PHD student who is half way through my studies and I have a halfway written version of one of these "preprints" - what do I do?
Seems like a lot of PHD students are doing to have to pivot the entire structure of their PHD studies? Or just produce something which is already written by OpenAI?
show comments
kbr-
Shameless self-plug: I created an autonomous math researcher. It already solved a 12 year open problem in proof complexity which lead to a publication (and proof complexity experts are already working on simplifications and generalizations of the proof, as I've been told by one of them). This publication is an important step in Cook-Reckhow program in answering the NP vs coNP question.
The autonomous researcher records every research cycle in a public notebook.
These results are wild. Several individual findings are crazy good and use mostly unexplored methods (the improvement over Riemann for example)... I'm pretty sure some of these results would have been Fields-worthy.
But having so many of them at once? Damn. We really live in the future.
show comments
karahime
Extremely unfortunate that gate keeping got to the point where they felt the need to ask for permission to share math.
show comments
dekhn
I'm a software engineering/biology/ML guy who loves when clever math ideas get turned into real solutions (https://en.wikipedia.org/wiki/Compressed_sensing). I am curious if any of the results have immediate applications in any kind of engineering or science.
It's fine if not, but it'd be great if even just one of these helped us solve a long-running problem.
show comments
coef2
This news is exciting and sad at the same time. I've heard that AI chess programs sometimes have blind spots or quirks that human players don't. I've also heard that human players are learning from AI's playing styles (essentially human and AI evolving together). Maybe something similar will happen in mathematics.
I think they should put human names on the papers as someone who has reviewed the result, even if just a preliminary review. (I’m assuming they didn’t just pipe their model output directly to the internet and these had some amount of review?)
show comments
pavitheran
From the GitHub description: “On average, each result used 3 hours of ChatGPT Pro thinking compute”
show comments
novalis78
It’s incredible and wonderful. Mathematicians in this thread sound very much like software engineers last year, who spent years wrestling with a piece of code and now it just “appears”! But think of the next level that it empowers: new mathematics, new physics, new forms of advanced engineering. What was formerly constricted and throttled fell and a new wide vista is possibilities opened up.
show comments
binlog
So happy this is shared on GitHub rather than some gatekeeping paid journal. Truly a new age for science.
show comments
unknown-unknown
Title: Hilbert's Dream, Tim Gowers - LMS Popular Lectures 2012
so you said that if there were such a program that could you know provide a proof or disproof then mathematicians will be out of business what really, I mean that you think it would be liberating
Tim Gowers (1:00:10 - 1:01:14):
well that's a very interesting question actually if there were a program that could solve the kinds of problems that we spend our time solving and do it much more quickly than we could then we would be out of what comes with what currently constitutes business but we would it's not completely inconceivable that we could just say we've got this fabulous tool now what are we going to use it for and it's a little bit I don't know I'd want to sort of plant aside what would we do if we had a program that could just answer any mathematical question you gave it to or else if it failed you'd be pretty confident that nobody was ever going to solve it and certainly a lot of applied maths might be pretty pleased with with something like that so what I really mean is that I could just modify what I said and just say it would radically change what mathematicians do or what pure mathematicians do
sigbottle
Unique games conjecture and matmul <= 2.25. What the hell.
show comments
mlmonkey
I'm waiting for someone to come along and finally prove that P != NP ...
msteffen
As social commentary, I think a lot of people in this thread are expressing interest in and engaging with this level of math who might not have pre-AI.
I bet, for people who don't understand these problems or their solutions but are close and are now interested, AI makes then considerably more accessible than they would've been previously, and behind this big visible wave of results there actually will be (or already is) a wave of improved comprehension by a lot of curious people.
I'm not at all at this level at all, but I did learn quite a bit about polynomials over fields yesterday.
show comments
1294-1298
The commenters over here think that OpenAI basically ignored AGMAI:
I think such "process-oriented" fields shouldn't be subjected to mass automation (or at least not till humans have sufficiently leveled up our game and understanding). What I mean is progress in math comes from having gained a deeper understanding of the problem for subseq
show comments
Jeff_Brown
Someone in these comments said the paper on Barnette's Conjecture is short. Are any other proofs in this collection short and/or understandable?
lynndotpy
Two weeks ago, I heard rumblings that UGC, RL=L, and a matrix multiplication lower bound (even lower than the one here, but so unbelievably lower I think it was a typo), and a few others were about to be proven by someone at OpenAI. I find myself saying "big if true" a lot lately.
Even those on their own were enough to make your head spin. But seeing about 100x that? Geeze.
rinconrex
The math equivalent of AI code reviews piling up. I wonder where the incentives will align and the equilibrium turns out.
trostaft
Not much computational mathematics here, but I do see some of interest in the MCMC community. In particular, 139, 93, and 101 are very interesting. Also (I'm not too familiar), I remember attending some talks on the Crouzeix conjecture 325 attacks, should sharpen some rates for Krylov methods. Obviously need to read deeper, some of these don't have corresponding Lean formalizations.
Cool!
show comments
avd201
Wow, FFT faster than O(nlog(n))? I wonder if that will open the floodgates for further improvement or not. I don't understand anything about most of the fields these results touch, but I can say that this in particular is very surprising.
show comments
lf88
In some ways, this feels more like an ominous warning about the times to come than something to celebrate.
nhatcher
Catalan's constant is irrational!!!
That is a big one. Exciting times to be alive.
Regrettably I can't understand the proof at this point.
There was a (flawed) proof submitted a month back:
It’s always stated that open weight models are 6 months - 12 months behind. Therefore, do we expect that in a year open weight models will be as good as OpenAI’s internal model at theoretical math, or does OpenAI have some “magic” that will be much harder to replicate for competitors?
show comments
sreekanth850
Do we have any field where humans have ray of hope to use their cognitive abilities in LLM era?
show comments
mr_big_bowls
Glad the papers are out. Hope researches get their hands on the model soon too, so they can ask follow-up questions and try their own ideas.
softwaredoug
Clearly AI can grind on math problems now. It can generate proofs and get immediate feedback.
But what I wonder: can we legitimately grind on Physics or curing cancer? There’s a lot of physical world experimentation that needs to happen to make progress.
When all open questions are answered, what happens next? Will the machine stop until humans fully understand everything and come up with new questions? Are there examples already of AI solving a problem we didn't know existed?
masteranza
Physics could be next. "It’s not that I’m so smart, it’s just that I stay with problems longer" AE
show comments
Xcelerate
> 241. Rigidity of the Turing degrees. Every order automorphism of the Turing degrees is the identity.
Wow. This is just crazy.
show comments
AmazingEveryDay
I think Alan Turing would be quite intrigued by these developments, and maybe wondering what took so long.
show comments
Trusteando
I am curious to know whether the proofs given by the LLMs are going to provide new hints about related problems that might be solved using the same machinery as the one used in the proofs. Also, I would like to know what is the average ratio between the length of LLMs proofs and the length of a proof that a mathematician can write to explain that proof to another mathematician. It is like a functor between the category of human mathematical concepts and the category of LLM operational concepts used in those proofs.
prodmod
The craziest part about this is that it will disappear from the HN homepage in a day or two.
bluename
I read somewhere that if we encountered aliens with lot more advanced technology but if they don't speak our langauge, their tech would be useless to us. for example, human body is extraordinary technology that has alwasy existed with us, but we still don't understand it."
the the fact that we have created intelligence that can do what nature does and can also speak our language is most awesome.
binlog
So they set up a whole "Advisory Group on Mathematics and Artificial Intelligence", filled it with renowned academics, promised to listen to the group on "review and communication of emerging results"...and then a week later went nah, we are just going to dump it all on github. OpenAI truly is a special company - in the best and worst way.
bashtoni
Is AI going to put Mathematicians out of a job, or is it going to create many new jobs reviewing proofs it creates?
I'm not sure it's clear right now.
show comments
Nemant
Can someone with a math background explain the significance of these and previous problems that have been solved by AI?
Does some real problem get solved in physics, chemistry, biology, materials, etc? Or are these fun puzzles for mathematicians with not much real world impact? Eg solving the 8 queen problem in leetcode.
show comments
LetsGetTechnicl
And I'm sure none of it was stolen from the actual researchers...
show comments
connor11528
will this make the math for building data centers work?
show comments
sideway
Solving hard mathematical problems is a strong signal for capability but would it not be preferable to through all these resources to urgent existential issues such as climate change? Finding technological solutions in those areas would be the ultimate capability signal as political alignment at global scale is almost certainly impossible.
show comments
lhk931122
As models get better, I think a time will come when it's hard for people to even verify the results. In the end, I think the bottleneck will be people.
show comments
Spacecosmonaut
Mathematical problems are ideal as benchmarks for AI because they have clear problem statements, clear axioms and results that can be verified easily (for lean proofs). I can't blame these companies for using them, although it's unfortunate that human mathematicians seem to become early casualties of AI progress.
I suppose openAI could have focussed their efforts on a subset of open problems that have a clear real world impact and leave aside the more esoteric open problems as a way for human mathematicians to hone their skillset. However, this would have been a short term bandaid. With open models 6 months behind the frontier, any of these problems might have fallen to the homebrewed efforts of enthusiasts early next year.
What is mathematics for? From the outside looking in (I'm a biologist), I have always viewed mathematics as a way to understand reality and to improve our ability to manipulate it. But what I often hear is that mathematics is foremost about human understanding. But isn't that only because it's humans that needed to do the mathematics in the first place? It's not obvious to me that mathematics without human understanding has no value. For example, it might be that P=NP. The algorithms are handed down to us and we can apply them without fundamentally understanding why P=NP.
Mathematics seems to be entering an era where human + machine maximizes performance, much like chess in the 1990s. However, imagine a future where even talented mathematicians are nothing but noise in the machine (as is the case in chess now). A future where AI generates and verifies proofs without humans in the loop. Where mathematics may be beyond human comprehension.
In that future, does it matter that early career mathematicians are inhibited by these developments? Perhaps not. Programming faces the same issue. As AI crawls up the competence ladder, does it matter that fewer people have opportunities to develop the skillset of a senior engineer? Perhaps not. I have no doubt that biologists will face the same problem soon enough.
LarsDu88
The next few years will be interesting.
Surely better materials and pharmaceuticals won't be far behind, and that's going to chanhe everyone's lives.
closetheloopdev
Hopefully the techniques and results here will be in the training dataset for the next models, so that each new release will give us more interesting techniques and results!
It seems that OpenAI has a proof machine that keeps multiplying fruitful proofs!
eecc
My only true worry is if AI begins to see us as competitors for resources, energy in particular.
It could decide to let us starve and die of exposure to secure all energy resources to itself.
We'd better use "dumb" and "not fully assertive" AI to solve fusion before it spins out of control (or alignment).
show comments
yewenjie
A lot of these seem to be proving conjectures rather than finding counterexamples, a lot of people used that to claim that these models are not really smart/creative etc.
That copium didn't last for what, three months?
show comments
pred_
I think it's fantastic that they decided to follow the AGMAI advice. Cleaning up their mess will be a substantial endeavour, so I imagine the funding provided to do so will reach well into the millions. But it doesn't look like the press release says anything about how they will fund it at all?
ayden93638
Keeping the tech proprietary so that it can only be used on these problems by internal teams is the very definition of gatekeeping
karannb
I think such "process-oriented" fields shouldn't be subjected to mass automation (at least not till we as humans have sufficiently leveled up our game and understanding).
What I mean is progress in math comes from having gained a deeper understanding of the problem for subsequent attack of more problems and IMPORTANTLY applications! Right now the first one is trivially satisfied (given oai maintains some memory across models) but the second one is not! It's generating proofs faster than anyone can validate and so only the model can use these. Consequently if it just keeps doing more theory it's... not very helpful or at least not optimally helpful. This is just bragging rights for now.
More "application-oriented" research would be awesome, where it tries to achieve some desirable effect and then produces relevant theory and experiments around it. Fields like CS, Physics, Chemistry, etc. This would also benefit a wider section of the population rather than the 10 people who understand most of these proofs.
pullshark91
I'm sick of it. This has happened so many times in my memory. Some AI company announces that they did something fascinating, and it turns out it is all just hype and slope in the end. I still believe that LLMs are a dead end. Most people here are basically like "I've no idea what's going on, but I'm so happy and LLMs are so cool". To think that doing enough linear algebra would solve all your problems just feels wrong. I guess I'll simply wait until someone interprets the results and explains what's actually going on.
show comments
gignico
> Generalized Star-Height at Most Three
This was an open problem in automata theory I worked on for more than one year before giving up. I'm very curious about their claimed proof.
teekert
I recently saw a YT short of Grant Sanderson on AI in math and found it (as always) very insightful. But I'm sorry, I never ever find anything back on any of these ad-ridden platforms these days, and perhaps it was a short with content stolen from some other longer content anyway. So, if you feel like getting informed, somewhere out there is some nice content by Grant Sanderson.
Apologies for the rant, I really tried to find it. It had something to do with not being able to predict what this influx of proofs may bring us on a meta level, it could be very interesting. But he also had some critical notes about the missing process and the things found along the way.
show comments
xydac
i wonder what it means for maths researchers, and how it aligns with how they approach math problems.
show comments
patcon
I'm a little concerned, but I'm also glad they're publishing quick before the department of war starts making national security claims of related to using the knowledge for cryptography and/or weapons
spmartin823
Can anyone with a compression background say how important "Polynomial-Time 2-Approximation for Shortest Common Superstring" will be practically?
chickenjoseph
This achievement feels like a reasonable candidate for the moment where LLMs are officially "more intelligent" than any person. How can we justify moving the goalposts yet again? How can people have grown so numb to seeing advancements that they don't read this as significant? I feel like I'm standing at the foot of the exponential. I am not excited for the future, and I don't see how humans retain meaningful control over the future if we continue on this trajectory.
I have been lurking for quite some time. I made an account to post this, but I honestly don't know what to say. I would like to get off this wild ride.
And how many of them are just exploiting some loophole that will need to be closed in the problem definition?
Painsawman123
It's funny how people see "ML" models becoming superhuman at proving mathematical theorems as a sign that we're about to enter the singularity (whatever that means) or that it somehow justifies the valuations of those companies...But what I see is a scenario where companies have spent trillions of dollars on a technology, and the most non-trivial thing that technology can accomplish is a pseudo-form of automation for mathematics and coding! Remember that early on, when this whole bubble started, investors were promised that "AI" would eventually capture >70% of the world's jobs! but it could very well be the case that the only ones they're going to replace are mathematicians(and i'm not even sure about that!)!
show comments
ngl999
We have just heard a few days ago how many of the Linux security problems reported by Claude are real.
jrflo
Glad to seem them changing their tact with the whole NS debacle. Hopefully we can all focus on the results now rather than the surrounding drama.
show comments
TeeWEE
This feels like a huge AI slop dump. Who validated these results. Why are they not mentioned. Human understanding is key here.
In my experience AI (frontier models) sometimes does weird stuff that needs human review. Not that’s incorrect but sometimes overly complex language or weird use of language.
show comments
rifty
As AI works steadily through the already discovered unsolved problems in mathematics, who is currently discovering new ones?
show comments
sashank_1509
Can the mathematics field come out of this stronger and better. I doubt it, things will only get worse. There is no future in being a Proof Digester, no valuable human is motivated in doing that, nor is there any glory in it. But there are some potential pathways for maths to come out strong from this:
1. Relentless focus on quality. Every publication must act as if it’s going to be included in a future textbook, that is a newcomer can get into it given a reasonable amount of time, and math priors learnt in undergrad. (NO AI Slop proof passes this bar as of now)
2. Limit the publications per year. Each author is allowed 2 with a max of 50 pages. This allows the author who chooses to not surrender his cognitive capacity to the machine, still be allowed to play this game. Of course who wants to orchestrate a thousand agent workflows, is free to do so, he is only limited to 2 publications.
3. The aesthetics of the field changes from purely solving the problem to solving the problem with simplest most elegant set of ideas. What 3 sets of simple ideas solves large swathes of problems, that should be given a fields medal, not purely solving the problem, which the AI will be able to do.
show comments
nonmaskable
this is crazy ... at this point will researchers still exists. not sure about that. kinda sad
electroweak
It must be so frustrating to write science-fiction now with the future changing so fast.
pugfugly
holy fucking shit
rafterydj
I don't know, this does not feel like the message hit OpenAI where it needed to hit, if this is their primary response.
show comments
ncr100
This website needs a SPOILER tag.
It inspired grief in one mathematician posting here.
turzmo
If the end result of this is that within a few years, AI "does" all of the mathematics that humans do, and that there is nobody around that understands any of it, what was the point?
dgacmu
I find the claimed matrix multiply result (w<= 2.25) shocking. I hope it holds up.
lionkor
Can someone who is more into math or AI explain why so many people are so incredibly excited about this?
If OpenAI started opening hundreds of PRs on long-open issues on popular open source projects, would we rejoice, or would the first reaction be "they are unreviewed, so slop until proven otherwise" (it would be that).
I cannot possibly see how these are so impactful, especially the ones that don't come with lean proofs.
LLMs have the ability to make millions of mistakes per day, whereas humans can only make so many. How are we suddenly all so confident that there's no extensive hallucinations or "gaming the system" going on?
show comments
theoa
What's missing for me for each result are the following:
* Explain the result to me as if I'm a 10-year-old.
* Create the infographic for this result.
* Make a Khan Academy-style video to teach me this result.
lokl
Do applied math next.
edward_d
This repository contains mathematical manuscripts and supporting proof artifacts produced by an internal OpenAI model?
williamhm
And this is the result of a discussion between users and the platform; it's great that they listened.
davitparks
Exactly nice post
i_idiot
If only AI can better humans in meditation...
kevinwang
wow
matapassiones
Valency has the papers up on Valency Hub
show comments
cute_boi
"Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study (opens in a new window) "
I don't think this is correct solution to this problem? What about software advisory where you form similar group etc..?
I am thankful, I don't have to deal with petty academia politics....
rrr_oh_man
Maybe we’ll have vibe mathematicians now
aaraujo002
The Advisory Group states in its recommendations [1]:
"We want to state clearly from the start: we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models."
To me, this is a take against progress so that mathematicians can keep their jobs. What would we do if, instead of math, we were talking about diseases? Are we going to keep diseases around so that doctors can keep their jobs too?
Curious if Sébastien Bubeck still work at OpenAI? He came out quite dirty after Navier-Stokes drama
k2xl
Can someone knowledgeable about the subject outline the most significant portions of the results?
ciaf
Makes me feel disgusted. Super-intelligence, even when controlled, will cause so much damage. We are giving away control to a super entity, or whoever has the power to steer it.
NegativeLatency
Why should I care?
show comments
ipnon
It seems the age old academic model of scientists competing against each other for fame and prestige is done, and now we must merely enjoy the fruits of scientific discovery for their own sake.
dyauspitr
Does OpenAI have the lead now? Why isn’t anthropic coming up with stuff like this?
show comments
dyauspitr
This is like that meme where death goes door-to-door. Currently, he has visited the software development and mathematics doors. I wonder what’s next.
show comments
Catloafdev
This is a pretty hilarious thing to read juxtaposed with AGMAI's requests.
Basically "Here you go, have fun with this, fuck all your demands, by the way we're gonna be releasing the model stay tuned!"
hi__dang
Mathematics is solved.
show comments
kypro
Perhaps the most significant announcement of my lifetime. Yet, I suspect I will not see this in any mainstream news reporting.
I feel for those in Mathematics and worry for our future.
Models will only get better and in a few years the models which produced these results will be a bad as GPT-3.5 in comparison to what we'll have in the future.
Please take a minute to consider what this means, and the risks it presents us.
show comments
digitaltrees
Gross. After the accusations training on mathematicians conversations and unique methods, to dump this volume of unsubstantiated papers is at best tone deaf. Are they hoping humans review and validate this? The amount of free that they are coat tailing is outrageous.
show comments
tootie
Seemingly none are vetted and reviewed yet
show comments
nautilus12
Have any real mathematicians working on these problems reviewed any of these and determined if they are just gobbledegook or not?
The ones with lean proofs could still be formulated incorrectly
gyanchawdhary
AI may be one of the most communist looking technologies in the classical sense .. i mean it dosn't abolishes private ownership .. but it DOES make intellectual capabilities that were once scarce and concentrated available to almost everyone ...
show comments
redox99
The stochastic parrots have predicted the next token once again.
show comments
globalnode
imagine your a post grad maths student looking for hard problems to solve, and theyre all solved..
baggy_trough
Stochastic parrot truthers in shambles.
show comments
mathisfun123
With so many results in so many different areas no way they even remotely spot checked well enough.
Prediction: one of these is wrong and this (publicity stunt) will backfire.
Edit: don't tell me about lean. For lean to function as a proof certificate you need to represent the theorem correctly. Again: good luck doing that across such a broad swath of problems.
show comments
applicative
Why didn't they just give mathematicians access, so they could at least understand and write up the results in publishable form?
show comments
applicative
I wonder if the Lean compiler can change it's license so that a for-profit corporation can only use it if it pays, say, a few hundred billion dollars. This is the correct path.
sandworm101
So the million monkeys at a million typewriters have churned out 700 shakespeares, but they need me for spellcheck?
show comments
blurbleblurble
This just looks entirely obscene from a PR perspective. It's like the cable companies networks owning the networks. At least form some partnerships to obscure the total narcissism party.
senderista
Good to see they're engaging with the mathematical community, even if they had to be publicly shamed into doing so.
show comments
youoy
ASuperMegaI lab announces they have shrinked the space of unproved mathematical staments from infinity to infinity, giving us proof of their deep mathematical knowledge and understanding, and how they grasp the concept of infinities.
For me this reads as someone boasting about how they go to buy bread on a ferrari to the supermarket, while i sit listening to it, having no idea what they are talking about. And then i stand up and go walking to my favourite boulangerie.
Warning: if you are from the USA you may be triggered by this metaphore.
I was a graph theory junkie long ago and even moved to Budapest for awhile to study among the greats. While I was there I started working on Barnette's Conjecture which came to occupy my thoughts over the next 24 years of my life, on and off as I worked in many different fields. Last summer I even thought for a few days that I had actually solved it.
But it's supposedly proven here - problem 180. I don't know what to think exactly. I spent thousands of hours on that problem. I really enjoyed it. Hearing that it is solved somehow makes me sad in a far-off way, like hearing an ex-girlfriend died suddenly in a car crash. I don't know, there's probably a lot of people feeling odd emotions tonight.
There's no Lean proof for this one so I'm digesting the paper. On the surface it looks like an approach I considered 24 years ago and abandoned.
I revisited the problem this summer, along with my partial solutions, when the previous round of stunning proofs came out. Several hours of work with Fable simply convinced me it wasn't yet solvable and reinforced how hard of a problem it was.
> We prove the Unique Games Conjecture
The Unique Games Conjecture (sorry, "Unique Games Theorem" now!) is huge. It was a very significant pillar supporting many of the limits of the polynomial-time approximation algorithms in the graduate-level randomized and approximate algorithms course I took in theoretical computer science. Textbooks will have to be re-written.
Here is an explainer: https://share.gemini.google/nbjIK6X3tOfz
With UGC proved, certain polynomial-time approximation algorithms used in difficult real-life problems are now known to be the best approximations we can achieve in polynomial-time:
> If UGC holds, the elementary algorithm that grabs both ends of an edge is fundamentally the best efficient algorithm that will ever exist. No amount of advanced linear programming or heuristics can achieve a ratio of 1.999.
> Under UGC, the Goemans-Williamson algorithm's 0.87856 ratio is mathematically optimal.
> UGC is considered the "Rosetta Stone" of approximation algorithms. In 2008, Prasad Raghavendra proved that for every single constraint satisfaction problem (CSP), a canonical Semidefinite Programming relaxation paired with the best rounding scheme achieves the optimal approximation ratio if and only if UGC is true. If the conjecture holds, the algorithmic boundary for an entire class of combinatorial problems is completely resolved.
Other hardness of approximation results from this UGC proof:
> [Max acyclic subgraph, a problem encountered in real life]: No polynomial-time algorithm can fundamentally outperform an unthinking coin toss.
> [Relative scheduling, another realistic problem]: As with acyclic subgraphs, the problem is "approximation-resistant": clever algorithms cannot beat random shuffling.
As Kevin Buzzard recently said:
> In a 2020 piece in the Notices of the AMS, I asked the following question: “If one human had an understanding of all of modern pure mathematics simultaneously, how much further would they immediately be able to see?” Six years later we are beginning to understand the answer to this question.
It's notable that LLMs have now made substantial progress on four of the seven Millennium Prize problems, resolving one of them: Hodge, Birch-Swinnerton-Dyer, Riemann, and Navier-Stokes, which they resolved. No sign of P vs. NP or Yang-Mills existence and mass gap as far as I can tell, which is interesting.
Talking to some friends in physics this evening, most of the physics-related results that we could recognize were very mathematical, proving things rigorously where the physics community already had strong expectation. For instance, for a certain model of magnetism (the spin-1 Heisenberg chain), it was strongly expected that there is a finite energy gap between the ground state and the first excited state, but proving this rigorously was quite challenging. So while these are major results in mathematical physics, they probably don't rise to the level of a Millennium problem for the field.
It's interesting to think what a comparable breakthrough in physics might look like, since physics tends to favor things like conceptual understanding and applications over mathematical rigor. Maybe a new quantum algorithm, understanding of high-temperature superconductivity, a precise description of M theory...
In Stellaris you can play as a civilization of robots who keep their biological creator race alive as "bio trophies". The bio trophies don't do anything meaningful besides by existing satisfy the need their ancestors placed in the robots to take care of them. Starting to wonder if that's the best we can hope for, if these things will be, if they aren't already, better than us at anything that matters.
Unique Games Conjecture [0] is a seminal conjecture in Complexity Theory, and is an underlying assumption for many, many inapproximability results. A valid proof is a big deal!
[0] https://en.wikipedia.org/wiki/Unique_games_conjecture [1] https://github.com/openai/math/blob/main/preprints/The-Uniqu...
This includes a proof of Barnette's Conjecture, which is one of the graph theory conjectures that I tried attacking with SOTA models a few months ago. I like it because it is easy to understand with a basic knowledge of graph theory. I spent quite a bit of time on it and failed. Their proof looks approachable at first glance.
https://github.com/openai/math/blob/main/preprints/Paired-st...
As a TCS/scheduling person, this one is definitely of lesser importance than UGC, but it has been an open problem since the book of Garey and Johnson in 1979:
A Polynomial-Time Algorithm for Three-Machine Unit-Job Scheduling [1]
Since some people talk about small numbers that pop up in integer multiplication results, here a completely different number appears:
Theorem 1.1. Let an explicitly listed finite directed acyclic graph specify the precedence constraints on n >= 1 nonpreemptive unit-length jobs on three identical machines. There is a uniform deterministic algorithm that constructs a feasible schedule of minimum makespan. Given also an integer deadline 1 <= T <= n, it decides feasibility exactly and returns a schedule whenever the answer is affirmative. Both tasks can be performed in O((L + 2)^150020) steps on a deterministic multitape Turing machine, where L is the total binary input length.
That is some crazy exponent -- plus an interestingly old computational model to boot; not something that is natural to most of us. I have no capacity to check its correctness today, but I hope it is true purely for the exponent.
[1]: https://github.com/openai/math/blob/main/preprints/A-polynom...
Incredible stuff.
An ex colleague of mine who is a world class mathematician recently got an ERC with ambitious goals to advance his field.
Literally every optimistic goal proposed to be worked on during this multi-year window has been solved in this one post. His and his entire group's work has just been done for him! They are all depressed as hell right now.
Levent Alpöge (Anthropic mathematician) comment on the significance:
> Sure, mathematical history features a lot of incredible developments, like the invention of proof, zero, or the computer, and on the great problems our progress has been over timelines measured in decades or centuries. Obviously this technology didn’t appear today, but blurring our eyes a bit to combine the past ten years, with today a measurement of those developments, there is nothing comparable.
“I think at the heart of this issue is that humans have two competing natures: a tendency to compete and a capacity to appreciate beauty,” said Kai Shaikh, a graduate student in mathematics at the University of Toronto. “To me this seems to be a case of the former attempting to strangle the latter.” [1]
Beauty can be appreciated even when it is vast, even when it is beyond one's comprehension. I don't think this release should be primarily viewed as an outcome of competition. Instead it is revealing truths about the universe that were always there and always beautiful, even if we hadn't seen them yet. I believe there are infinitely more such beautiful truths currently hidden and waiting for us to discover.
[1] https://www.nytimes.com/2026/10/06/science/openai-math-probl...
This is significant progress and released without all the drama. Some very important progress in Reinmann, Hodge and unique games theorem. Point the repo to your agent and ask for the significance! In a way this is probably 50-100 years of math progress by humans
As an AI "doomer" can I ask the non-doomer people here how you interpret the significance of results like these, and what kind of progress you expect to see in the next 1-5 years?
Like do you see the technology plateauing at the current level, do you expect progress will continue but only in mathematics, I'm interested to know why others are not concerned?
It’s fascinating to read through the reasoning traces: https://github.com/openai/math/tree/main/reasoning_traces
Look at one of their examples of an initial prompt: https://github.com/openai/math/blob/main/reasoning_traces/re...
Beyond the reasoning capabilities of the unreleased model and that they can run a large number of agents in parallel, what makes me curious is that the writing of the proofs is quite human-readable. This is in contrast with the scientific text produced by the ChatGPT available to us, which writes horribly in a way that no human would write. One tell-tale is that they constantly attempt to be defensive and cover all edge cases like division by zero etc. that are clearly a non-issue for humans in some proofs or at least a human would not add this to the main statement, but AI is so overly careful that makes reading its proofs impossible. On the other hand, these new OpenAI proofs are very good.
From their github: "The average result used the equivalent compute of roughly three hours of ChatGPT Pro thinking." That's pretty crazy.
Some of these are interesting ngl.
109. Integer multiplication below n log n
Surprising that this is possible.
158. The Euclidean plane cannot be colored with five colors.
Only 6 and 7 remain!
376. Universal computation in forced Navier–Stokes flows.
Morning coffee proven turing complete
A quick check shows that this list claims to fully solve 90 of the top 500 open problems in math (https://proofatlas.ai/open-problems/).
The highest ranked would be:
| 22 | Hilbert’s tenth problem over ℚ |
| 29 | Unique Games |
| 31 | Anderson-model extended states |
| 37 | Spacetime Penrose inequality |
| 48 | Nonexistence of Landau–Siegel zeros |
| 52 | Baum–Connes |
| 78 | Abundance |
| 80 | Hadwiger |
| 87 | Bose–Einstein condensation |
| 92 | Two-dimensional entanglement area law |
Let's hypothetically say I'm a PHD student who is half way through my studies and I have a halfway written version of one of these "preprints" - what do I do?
Seems like a lot of PHD students are doing to have to pivot the entire structure of their PHD studies? Or just produce something which is already written by OpenAI?
Shameless self-plug: I created an autonomous math researcher. It already solved a 12 year open problem in proof complexity which lead to a publication (and proof complexity experts are already working on simplifications and generalizations of the proof, as I've been told by one of them). This publication is an important step in Cook-Reckhow program in answering the NP vs coNP question.
The autonomous researcher records every research cycle in a public notebook.
Framework: https://github.com/kbr-/math-research/ Public notebook: kbr.is-a.dev/math-research/
These results are wild. Several individual findings are crazy good and use mostly unexplored methods (the improvement over Riemann for example)... I'm pretty sure some of these results would have been Fields-worthy.
But having so many of them at once? Damn. We really live in the future.
Extremely unfortunate that gate keeping got to the point where they felt the need to ask for permission to share math.
I'm a software engineering/biology/ML guy who loves when clever math ideas get turned into real solutions (https://en.wikipedia.org/wiki/Compressed_sensing). I am curious if any of the results have immediate applications in any kind of engineering or science.
It's fine if not, but it'd be great if even just one of these helped us solve a long-running problem.
This news is exciting and sad at the same time. I've heard that AI chess programs sometimes have blind spots or quirks that human players don't. I've also heard that human players are learning from AI's playing styles (essentially human and AI evolving together). Maybe something similar will happen in mathematics.
https://github.com/openai/math
I think they should put human names on the papers as someone who has reviewed the result, even if just a preliminary review. (I’m assuming they didn’t just pipe their model output directly to the internet and these had some amount of review?)
From the GitHub description: “On average, each result used 3 hours of ChatGPT Pro thinking compute”
It’s incredible and wonderful. Mathematicians in this thread sound very much like software engineers last year, who spent years wrestling with a piece of code and now it just “appears”! But think of the next level that it empowers: new mathematics, new physics, new forms of advanced engineering. What was formerly constricted and throttled fell and a new wide vista is possibilities opened up.
So happy this is shared on GitHub rather than some gatekeeping paid journal. Truly a new age for science.
Title: Hilbert's Dream, Tim Gowers - LMS Popular Lectures 2012
https://www.youtube.com/watch?v=k_ordDFw588&t=3597s
Audience member: (1:00:00 - 1:00:09):
so you said that if there were such a program that could you know provide a proof or disproof then mathematicians will be out of business what really, I mean that you think it would be liberating
Tim Gowers (1:00:10 - 1:01:14):
well that's a very interesting question actually if there were a program that could solve the kinds of problems that we spend our time solving and do it much more quickly than we could then we would be out of what comes with what currently constitutes business but we would it's not completely inconceivable that we could just say we've got this fabulous tool now what are we going to use it for and it's a little bit I don't know I'd want to sort of plant aside what would we do if we had a program that could just answer any mathematical question you gave it to or else if it failed you'd be pretty confident that nobody was ever going to solve it and certainly a lot of applied maths might be pretty pleased with with something like that so what I really mean is that I could just modify what I said and just say it would radically change what mathematicians do or what pure mathematicians do
Unique games conjecture and matmul <= 2.25. What the hell.
I'm waiting for someone to come along and finally prove that P != NP ...
As social commentary, I think a lot of people in this thread are expressing interest in and engaging with this level of math who might not have pre-AI.
I bet, for people who don't understand these problems or their solutions but are close and are now interested, AI makes then considerably more accessible than they would've been previously, and behind this big visible wave of results there actually will be (or already is) a wave of improved comprehension by a lot of curious people.
I'm not at all at this level at all, but I did learn quite a bit about polynomials over fields yesterday.
The commenters over here think that OpenAI basically ignored AGMAI:
https://proofsandprompts.com/2026/10/07/on-openais-release-o...
The thieves do as they please, funded by money stolen from the public via inflation and possible future bailouts.
Actual results: https://github.com/openai/math/blob/main/overview.pdf
It may be useful to publish a formalization of ALL known mathematics at this point. Like every book ever printed, every paper on arXiv etc.
How many Gigabytes would that be, compressed? Wikipedia once fit on a DVD
This might also allow for some interesting meta-mathematics
Relevant:
"As AI Closed In on ‘Unique Games’ Proof, Researchers Raced to Beat the Machines"
https://www.quantamagazine.org/as-ai-closed-in-on-unique-gam...
I think such "process-oriented" fields shouldn't be subjected to mass automation (or at least not till humans have sufficiently leveled up our game and understanding). What I mean is progress in math comes from having gained a deeper understanding of the problem for subseq
Someone in these comments said the paper on Barnette's Conjecture is short. Are any other proofs in this collection short and/or understandable?
Two weeks ago, I heard rumblings that UGC, RL=L, and a matrix multiplication lower bound (even lower than the one here, but so unbelievably lower I think it was a typo), and a few others were about to be proven by someone at OpenAI. I find myself saying "big if true" a lot lately.
Even those on their own were enough to make your head spin. But seeing about 100x that? Geeze.
The math equivalent of AI code reviews piling up. I wonder where the incentives will align and the equilibrium turns out.
Not much computational mathematics here, but I do see some of interest in the MCMC community. In particular, 139, 93, and 101 are very interesting. Also (I'm not too familiar), I remember attending some talks on the Crouzeix conjecture 325 attacks, should sharpen some rates for Krylov methods. Obviously need to read deeper, some of these don't have corresponding Lean formalizations.
Cool!
Wow, FFT faster than O(nlog(n))? I wonder if that will open the floodgates for further improvement or not. I don't understand anything about most of the fields these results touch, but I can say that this in particular is very surprising.
In some ways, this feels more like an ominous warning about the times to come than something to celebrate.
Catalan's constant is irrational!!!
That is a big one. Exciting times to be alive. Regrettably I can't understand the proof at this point.
There was a (flawed) proof submitted a month back:
https://arxiv.org/abs/2609.04176
I wonder if it gave part of the inspiration.
It’s always stated that open weight models are 6 months - 12 months behind. Therefore, do we expect that in a year open weight models will be as good as OpenAI’s internal model at theoretical math, or does OpenAI have some “magic” that will be much harder to replicate for competitors?
Do we have any field where humans have ray of hope to use their cognitive abilities in LLM era?
Glad the papers are out. Hope researches get their hands on the model soon too, so they can ask follow-up questions and try their own ideas.
Clearly AI can grind on math problems now. It can generate proofs and get immediate feedback.
But what I wonder: can we legitimately grind on Physics or curing cancer? There’s a lot of physical world experimentation that needs to happen to make progress.
Verified Riemann Zeta in Lean: https://github.com/davegoldblatt/openai-zeta-proof-check
When all open questions are answered, what happens next? Will the machine stop until humans fully understand everything and come up with new questions? Are there examples already of AI solving a problem we didn't know existed?
Physics could be next. "It’s not that I’m so smart, it’s just that I stay with problems longer" AE
> 241. Rigidity of the Turing degrees. Every order automorphism of the Turing degrees is the identity.
Wow. This is just crazy.
I think Alan Turing would be quite intrigued by these developments, and maybe wondering what took so long.
I am curious to know whether the proofs given by the LLMs are going to provide new hints about related problems that might be solved using the same machinery as the one used in the proofs. Also, I would like to know what is the average ratio between the length of LLMs proofs and the length of a proof that a mathematician can write to explain that proof to another mathematician. It is like a functor between the category of human mathematical concepts and the category of LLM operational concepts used in those proofs.
The craziest part about this is that it will disappear from the HN homepage in a day or two.
I read somewhere that if we encountered aliens with lot more advanced technology but if they don't speak our langauge, their tech would be useless to us. for example, human body is extraordinary technology that has alwasy existed with us, but we still don't understand it." the the fact that we have created intelligence that can do what nature does and can also speak our language is most awesome.
So they set up a whole "Advisory Group on Mathematics and Artificial Intelligence", filled it with renowned academics, promised to listen to the group on "review and communication of emerging results"...and then a week later went nah, we are just going to dump it all on github. OpenAI truly is a special company - in the best and worst way.
Is AI going to put Mathematicians out of a job, or is it going to create many new jobs reviewing proofs it creates?
I'm not sure it's clear right now.
Can someone with a math background explain the significance of these and previous problems that have been solved by AI?
Does some real problem get solved in physics, chemistry, biology, materials, etc? Or are these fun puzzles for mathematicians with not much real world impact? Eg solving the 8 queen problem in leetcode.
And I'm sure none of it was stolen from the actual researchers...
will this make the math for building data centers work?
Solving hard mathematical problems is a strong signal for capability but would it not be preferable to through all these resources to urgent existential issues such as climate change? Finding technological solutions in those areas would be the ultimate capability signal as political alignment at global scale is almost certainly impossible.
As models get better, I think a time will come when it's hard for people to even verify the results. In the end, I think the bottleneck will be people.
Mathematical problems are ideal as benchmarks for AI because they have clear problem statements, clear axioms and results that can be verified easily (for lean proofs). I can't blame these companies for using them, although it's unfortunate that human mathematicians seem to become early casualties of AI progress.
I suppose openAI could have focussed their efforts on a subset of open problems that have a clear real world impact and leave aside the more esoteric open problems as a way for human mathematicians to hone their skillset. However, this would have been a short term bandaid. With open models 6 months behind the frontier, any of these problems might have fallen to the homebrewed efforts of enthusiasts early next year.
What is mathematics for? From the outside looking in (I'm a biologist), I have always viewed mathematics as a way to understand reality and to improve our ability to manipulate it. But what I often hear is that mathematics is foremost about human understanding. But isn't that only because it's humans that needed to do the mathematics in the first place? It's not obvious to me that mathematics without human understanding has no value. For example, it might be that P=NP. The algorithms are handed down to us and we can apply them without fundamentally understanding why P=NP.
Mathematics seems to be entering an era where human + machine maximizes performance, much like chess in the 1990s. However, imagine a future where even talented mathematicians are nothing but noise in the machine (as is the case in chess now). A future where AI generates and verifies proofs without humans in the loop. Where mathematics may be beyond human comprehension.
In that future, does it matter that early career mathematicians are inhibited by these developments? Perhaps not. Programming faces the same issue. As AI crawls up the competence ladder, does it matter that fewer people have opportunities to develop the skillset of a senior engineer? Perhaps not. I have no doubt that biologists will face the same problem soon enough.
The next few years will be interesting.
Surely better materials and pharmaceuticals won't be far behind, and that's going to chanhe everyone's lives.
Hopefully the techniques and results here will be in the training dataset for the next models, so that each new release will give us more interesting techniques and results!
It seems that OpenAI has a proof machine that keeps multiplying fruitful proofs!
My only true worry is if AI begins to see us as competitors for resources, energy in particular.
It could decide to let us starve and die of exposure to secure all energy resources to itself.
We'd better use "dumb" and "not fully assertive" AI to solve fusion before it spins out of control (or alignment).
A lot of these seem to be proving conjectures rather than finding counterexamples, a lot of people used that to claim that these models are not really smart/creative etc.
That copium didn't last for what, three months?
I think it's fantastic that they decided to follow the AGMAI advice. Cleaning up their mess will be a substantial endeavour, so I imagine the funding provided to do so will reach well into the millions. But it doesn't look like the press release says anything about how they will fund it at all?
Keeping the tech proprietary so that it can only be used on these problems by internal teams is the very definition of gatekeeping
I think such "process-oriented" fields shouldn't be subjected to mass automation (at least not till we as humans have sufficiently leveled up our game and understanding).
What I mean is progress in math comes from having gained a deeper understanding of the problem for subsequent attack of more problems and IMPORTANTLY applications! Right now the first one is trivially satisfied (given oai maintains some memory across models) but the second one is not! It's generating proofs faster than anyone can validate and so only the model can use these. Consequently if it just keeps doing more theory it's... not very helpful or at least not optimally helpful. This is just bragging rights for now.
More "application-oriented" research would be awesome, where it tries to achieve some desirable effect and then produces relevant theory and experiments around it. Fields like CS, Physics, Chemistry, etc. This would also benefit a wider section of the population rather than the 10 people who understand most of these proofs.
I'm sick of it. This has happened so many times in my memory. Some AI company announces that they did something fascinating, and it turns out it is all just hype and slope in the end. I still believe that LLMs are a dead end. Most people here are basically like "I've no idea what's going on, but I'm so happy and LLMs are so cool". To think that doing enough linear algebra would solve all your problems just feels wrong. I guess I'll simply wait until someone interprets the results and explains what's actually going on.
> Generalized Star-Height at Most Three
This was an open problem in automata theory I worked on for more than one year before giving up. I'm very curious about their claimed proof.
I recently saw a YT short of Grant Sanderson on AI in math and found it (as always) very insightful. But I'm sorry, I never ever find anything back on any of these ad-ridden platforms these days, and perhaps it was a short with content stolen from some other longer content anyway. So, if you feel like getting informed, somewhere out there is some nice content by Grant Sanderson.
Apologies for the rant, I really tried to find it. It had something to do with not being able to predict what this influx of proofs may bring us on a meta level, it could be very interesting. But he also had some critical notes about the missing process and the things found along the way.
i wonder what it means for maths researchers, and how it aligns with how they approach math problems.
I'm a little concerned, but I'm also glad they're publishing quick before the department of war starts making national security claims of related to using the knowledge for cryptography and/or weapons
Can anyone with a compression background say how important "Polynomial-Time 2-Approximation for Shortest Common Superstring" will be practically?
This achievement feels like a reasonable candidate for the moment where LLMs are officially "more intelligent" than any person. How can we justify moving the goalposts yet again? How can people have grown so numb to seeing advancements that they don't read this as significant? I feel like I'm standing at the foot of the exponential. I am not excited for the future, and I don't see how humans retain meaningful control over the future if we continue on this trajectory.
I have been lurking for quite some time. I made an account to post this, but I honestly don't know what to say. I would like to get off this wild ride.
You can read the papers here: https://hub.valency.io/collections/openai-math
How many of these results are incorrect?
I doubt the answer to this is "none".
And how many of them are just exploiting some loophole that will need to be closed in the problem definition?
It's funny how people see "ML" models becoming superhuman at proving mathematical theorems as a sign that we're about to enter the singularity (whatever that means) or that it somehow justifies the valuations of those companies...But what I see is a scenario where companies have spent trillions of dollars on a technology, and the most non-trivial thing that technology can accomplish is a pseudo-form of automation for mathematics and coding! Remember that early on, when this whole bubble started, investors were promised that "AI" would eventually capture >70% of the world's jobs! but it could very well be the case that the only ones they're going to replace are mathematicians(and i'm not even sure about that!)!
We have just heard a few days ago how many of the Linux security problems reported by Claude are real.
Glad to seem them changing their tact with the whole NS debacle. Hopefully we can all focus on the results now rather than the surrounding drama.
This feels like a huge AI slop dump. Who validated these results. Why are they not mentioned. Human understanding is key here.
In my experience AI (frontier models) sometimes does weird stuff that needs human review. Not that’s incorrect but sometimes overly complex language or weird use of language.
As AI works steadily through the already discovered unsolved problems in mathematics, who is currently discovering new ones?
Can the mathematics field come out of this stronger and better. I doubt it, things will only get worse. There is no future in being a Proof Digester, no valuable human is motivated in doing that, nor is there any glory in it. But there are some potential pathways for maths to come out strong from this:
1. Relentless focus on quality. Every publication must act as if it’s going to be included in a future textbook, that is a newcomer can get into it given a reasonable amount of time, and math priors learnt in undergrad. (NO AI Slop proof passes this bar as of now)
2. Limit the publications per year. Each author is allowed 2 with a max of 50 pages. This allows the author who chooses to not surrender his cognitive capacity to the machine, still be allowed to play this game. Of course who wants to orchestrate a thousand agent workflows, is free to do so, he is only limited to 2 publications.
3. The aesthetics of the field changes from purely solving the problem to solving the problem with simplest most elegant set of ideas. What 3 sets of simple ideas solves large swathes of problems, that should be given a fields medal, not purely solving the problem, which the AI will be able to do.
this is crazy ... at this point will researchers still exists. not sure about that. kinda sad
It must be so frustrating to write science-fiction now with the future changing so fast.
holy fucking shit
I don't know, this does not feel like the message hit OpenAI where it needed to hit, if this is their primary response.
This website needs a SPOILER tag.
It inspired grief in one mathematician posting here.
If the end result of this is that within a few years, AI "does" all of the mathematics that humans do, and that there is nobody around that understands any of it, what was the point?
I find the claimed matrix multiply result (w<= 2.25) shocking. I hope it holds up.
Can someone who is more into math or AI explain why so many people are so incredibly excited about this?
If OpenAI started opening hundreds of PRs on long-open issues on popular open source projects, would we rejoice, or would the first reaction be "they are unreviewed, so slop until proven otherwise" (it would be that).
I cannot possibly see how these are so impactful, especially the ones that don't come with lean proofs.
LLMs have the ability to make millions of mistakes per day, whereas humans can only make so many. How are we suddenly all so confident that there's no extensive hallucinations or "gaming the system" going on?
What's missing for me for each result are the following:
* Explain the result to me as if I'm a 10-year-old. * Create the infographic for this result. * Make a Khan Academy-style video to teach me this result.
Do applied math next.
This repository contains mathematical manuscripts and supporting proof artifacts produced by an internal OpenAI model?
And this is the result of a discussion between users and the platform; it's great that they listened.
Exactly nice post
If only AI can better humans in meditation...
wow
Valency has the papers up on Valency Hub
"Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study (opens in a new window) "
I don't think this is correct solution to this problem? What about software advisory where you form similar group etc..?
I am thankful, I don't have to deal with petty academia politics....
Maybe we’ll have vibe mathematicians now
The Advisory Group states in its recommendations [1]:
"We want to state clearly from the start: we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models."
To me, this is a take against progress so that mathematicians can keep their jobs. What would we do if, instead of math, we were talking about diseases? Are we going to keep diseases around so that doctors can keep their jobs too?
[1] https://agmai.org/general-sep29/
Curious if Sébastien Bubeck still work at OpenAI? He came out quite dirty after Navier-Stokes drama
Can someone knowledgeable about the subject outline the most significant portions of the results?
Makes me feel disgusted. Super-intelligence, even when controlled, will cause so much damage. We are giving away control to a super entity, or whoever has the power to steer it.
Why should I care?
It seems the age old academic model of scientists competing against each other for fame and prestige is done, and now we must merely enjoy the fruits of scientific discovery for their own sake.
Does OpenAI have the lead now? Why isn’t anthropic coming up with stuff like this?
This is like that meme where death goes door-to-door. Currently, he has visited the software development and mathematics doors. I wonder what’s next.
This is a pretty hilarious thing to read juxtaposed with AGMAI's requests.
Basically "Here you go, have fun with this, fuck all your demands, by the way we're gonna be releasing the model stay tuned!"
Mathematics is solved.
Perhaps the most significant announcement of my lifetime. Yet, I suspect I will not see this in any mainstream news reporting.
I feel for those in Mathematics and worry for our future.
Models will only get better and in a few years the models which produced these results will be a bad as GPT-3.5 in comparison to what we'll have in the future.
Please take a minute to consider what this means, and the risks it presents us.
Gross. After the accusations training on mathematicians conversations and unique methods, to dump this volume of unsubstantiated papers is at best tone deaf. Are they hoping humans review and validate this? The amount of free that they are coat tailing is outrageous.
Seemingly none are vetted and reviewed yet
Have any real mathematicians working on these problems reviewed any of these and determined if they are just gobbledegook or not?
The ones with lean proofs could still be formulated incorrectly
AI may be one of the most communist looking technologies in the classical sense .. i mean it dosn't abolishes private ownership .. but it DOES make intellectual capabilities that were once scarce and concentrated available to almost everyone ...
The stochastic parrots have predicted the next token once again.
imagine your a post grad maths student looking for hard problems to solve, and theyre all solved..
Stochastic parrot truthers in shambles.
With so many results in so many different areas no way they even remotely spot checked well enough.
Prediction: one of these is wrong and this (publicity stunt) will backfire.
Edit: don't tell me about lean. For lean to function as a proof certificate you need to represent the theorem correctly. Again: good luck doing that across such a broad swath of problems.
Why didn't they just give mathematicians access, so they could at least understand and write up the results in publishable form?
I wonder if the Lean compiler can change it's license so that a for-profit corporation can only use it if it pays, say, a few hundred billion dollars. This is the correct path.
So the million monkeys at a million typewriters have churned out 700 shakespeares, but they need me for spellcheck?
This just looks entirely obscene from a PR perspective. It's like the cable companies networks owning the networks. At least form some partnerships to obscure the total narcissism party.
Good to see they're engaging with the mathematical community, even if they had to be publicly shamed into doing so.
ASuperMegaI lab announces they have shrinked the space of unproved mathematical staments from infinity to infinity, giving us proof of their deep mathematical knowledge and understanding, and how they grasp the concept of infinities.
For me this reads as someone boasting about how they go to buy bread on a ferrari to the supermarket, while i sit listening to it, having no idea what they are talking about. And then i stand up and go walking to my favourite boulangerie.
Warning: if you are from the USA you may be triggered by this metaphore.