Sharing AI progress in mathematics

1207 points1362 commentsa day ago
jboggan

I was a graph theory junkie long ago and even moved to Budapest for awhile to study among the greats. While I was there I started working on Barnette's Conjecture which came to occupy my thoughts over the next 24 years of my life, on and off as I worked in many different fields. Last summer I even thought for a few days that I had actually solved it.

But it's supposedly proven here - problem 180. I don't know what to think exactly. I spent thousands of hours on that problem. I really enjoyed it. Hearing that it is solved somehow makes me sad in a far-off way, like hearing an ex-girlfriend died suddenly in a car crash. I don't know, there's probably a lot of people feeling odd emotions tonight.

There's no Lean proof for this one so I'm digesting the paper. On the surface it looks like an approach I considered 24 years ago and abandoned.

I revisited the problem this summer, along with my partial solutions, when the previous round of stunning proofs came out. Several hours of work with Fable simply convinced me it wasn't yet solvable and reinforced how hard of a problem it was.

show comments
winfieldchen

> We prove the Unique Games Conjecture

The Unique Games Conjecture (sorry, "Unique Games Theorem" now!) is huge. It was a very significant pillar supporting many of the limits of the polynomial-time approximation algorithms in the graduate-level randomized and approximate algorithms course I took in theoretical computer science. Textbooks will have to be re-written.

Here is an explainer: https://share.gemini.google/nbjIK6X3tOfz

With UGC proved, certain polynomial-time approximation algorithms used in difficult real-life problems are now known to be the best approximations we can achieve in polynomial-time:

> If UGC holds, the elementary algorithm that grabs both ends of an edge is fundamentally the best efficient algorithm that will ever exist. No amount of advanced linear programming or heuristics can achieve a ratio of 1.999.

> Under UGC, the Goemans-Williamson algorithm's 0.87856 ratio is mathematically optimal.

> UGC is considered the "Rosetta Stone" of approximation algorithms. In 2008, Prasad Raghavendra proved that for every single constraint satisfaction problem (CSP), a canonical Semidefinite Programming relaxation paired with the best rounding scheme achieves the optimal approximation ratio if and only if UGC is true. If the conjecture holds, the algorithmic boundary for an entire class of combinatorial problems is completely resolved.

Other hardness of approximation results from this UGC proof:

> [Max acyclic subgraph, a problem encountered in real life]: No polynomial-time algorithm can fundamentally outperform an unthinking coin toss.

> [Relative scheduling, another realistic problem]: As with acyclic subgraphs, the problem is "approximation-resistant": clever algorithms cannot beat random shuffling.

show comments
xanderlewis

As Kevin Buzzard recently said:

> In a 2020 piece in the Notices of the AMS, I asked the following question: “If one human had an understanding of all of modern pure mathematics simultaneously, how much further would they immediately be able to see?” Six years later we are beginning to understand the answer to this question.

show comments
qnleigh

It's notable that LLMs have now made substantial progress on four of the seven Millennium Prize problems, resolving one of them: Hodge, Birch-Swinnerton-Dyer, Riemann, and Navier-Stokes, which they resolved. No sign of P vs. NP or Yang-Mills existence and mass gap as far as I can tell, which is interesting.

Talking to some friends in physics this evening, most of the physics-related results that we could recognize were very mathematical, proving things rigorously where the physics community already had strong expectation. For instance, for a certain model of magnetism (the spin-1 Heisenberg chain), it was strongly expected that there is a finite energy gap between the ground state and the first excited state, but proving this rigorously was quite challenging. So while these are major results in mathematical physics, they probably don't rise to the level of a Millennium problem for the field.

It's interesting to think what a comparable breakthrough in physics might look like, since physics tends to favor things like conceptual understanding and applications over mathematical rigor. Maybe a new quantum algorithm, understanding of high-temperature superconductivity, a precise description of M theory...

show comments
rcr-anti

In Stellaris you can play as a civilization of robots who keep their biological creator race alive as "bio trophies". The bio trophies don't do anything meaningful besides by existing satisfy the need their ancestors placed in the robots to take care of them. Starting to wonder if that's the best we can hope for, if these things will be, if they aren't already, better than us at anything that matters.

show comments
enoether

Unique Games Conjecture [0] is a seminal conjecture in Complexity Theory, and is an underlying assumption for many, many inapproximability results. A valid proof is a big deal!

[0] https://en.wikipedia.org/wiki/Unique_games_conjecture [1] https://github.com/openai/math/blob/main/preprints/The-Uniqu...

show comments
prideout

This includes a proof of Barnette's Conjecture, which is one of the graph theory conjectures that I tried attacking with SOTA models a few months ago. I like it because it is easy to understand with a basic knowledge of graph theory. I spent quite a bit of time on it and failed. Their proof looks approachable at first glance.

https://github.com/openai/math/blob/main/preprints/Paired-st...

show comments
NotOscarWilde

As a TCS/scheduling person, this one is definitely of lesser importance than UGC, but it has been an open problem since the book of Garey and Johnson in 1979:

A Polynomial-Time Algorithm for Three-Machine Unit-Job Scheduling [1]

Since some people talk about small numbers that pop up in integer multiplication results, here a completely different number appears:

Theorem 1.1. Let an explicitly listed finite directed acyclic graph specify the precedence constraints on n >= 1 nonpreemptive unit-length jobs on three identical machines. There is a uniform deterministic algorithm that constructs a feasible schedule of minimum makespan. Given also an integer deadline 1 <= T <= n, it decides feasibility exactly and returns a schedule whenever the answer is affirmative. Both tasks can be performed in O((L + 2)^150020) steps on a deterministic multitape Turing machine, where L is the total binary input length.

That is some crazy exponent -- plus an interestingly old computational model to boot; not something that is natural to most of us. I have no capacity to check its correctness today, but I hope it is true purely for the exponent.

[1]: https://github.com/openai/math/blob/main/preprints/A-polynom...

show comments
Turneyboy

Incredible stuff.

An ex colleague of mine who is a world class mathematician recently got an ERC with ambitious goals to advance his field.

Literally every optimistic goal proposed to be worked on during this multi-year window has been solved in this one post. His and his entire group's work has just been done for him! They are all depressed as hell right now.

show comments
schleck8

Levent Alpöge (Anthropic mathematician) comment on the significance:

> Sure, mathematical history features a lot of incredible developments, like the invention of proof, zero, or the computer, and on the great problems our progress has been over timelines measured in decades or centuries. Obviously this technology didn’t appear today, but blurring our eyes a bit to combine the past ten years, with today a measurement of those developments, there is nothing comparable.

show comments
bcatanzaro

“I think at the heart of this issue is that humans have two competing natures: a tendency to compete and a capacity to appreciate beauty,” said Kai Shaikh, a graduate student in mathematics at the University of Toronto. “To me this seems to be a case of the former attempting to strangle the latter.” [1]

Beauty can be appreciated even when it is vast, even when it is beyond one's comprehension. I don't think this release should be primarily viewed as an outcome of competition. Instead it is revealing truths about the universe that were always there and always beautiful, even if we hadn't seen them yet. I believe there are infinitely more such beautiful truths currently hidden and waiting for us to discover.

[1] https://www.nytimes.com/2026/10/06/science/openai-math-probl...

show comments
gizmodo59

This is significant progress and released without all the drama. Some very important progress in Reinmann, Hodge and unique games theorem. Point the repo to your agent and ask for the significance! In a way this is probably 50-100 years of math progress by humans

show comments
againstapples

As an AI "doomer" can I ask the non-doomer people here how you interpret the significance of results like these, and what kind of progress you expect to see in the next 1-5 years?

Like do you see the technology plateauing at the current level, do you expect progress will continue but only in mathematics, I'm interested to know why others are not concerned?

show comments
sebmellen

It’s fascinating to read through the reasoning traces: https://github.com/openai/math/tree/main/reasoning_traces

Look at one of their examples of an initial prompt: https://github.com/openai/math/blob/main/reasoning_traces/re...

show comments
spatalo

Beyond the reasoning capabilities of the unreleased model and that they can run a large number of agents in parallel, what makes me curious is that the writing of the proofs is quite human-readable. This is in contrast with the scientific text produced by the ChatGPT available to us, which writes horribly in a way that no human would write. One tell-tale is that they constantly attempt to be defensive and cover all edge cases like division by zero etc. that are clearly a non-issue for humans in some proofs or at least a human would not add this to the main statement, but AI is so overly careful that makes reading its proofs impossible. On the other hand, these new OpenAI proofs are very good.

foota

From their github: "The average result used the equivalent compute of roughly three hours of ChatGPT Pro thinking." That's pretty crazy.

show comments
kingstnap

Some of these are interesting ngl.

109. Integer multiplication below n log n

Surprising that this is possible.

158. The Euclidean plane cannot be colored with five colors.

Only 6 and 7 remain!

376. Universal computation in forced Navier–Stokes flows.

Morning coffee proven turing complete

show comments
zone411

A quick check shows that this list claims to fully solve 90 of the top 500 open problems in math (https://proofatlas.ai/open-problems/).

The highest ranked would be:

| 22 | Hilbert’s tenth problem over ℚ |

| 29 | Unique Games |

| 31 | Anderson-model extended states |

| 37 | Spacetime Penrose inequality |

| 48 | Nonexistence of Landau–Siegel zeros |

| 52 | Baum–Connes |

| 78 | Abundance |

| 80 | Hadwiger |

| 87 | Bose–Einstein condensation |

| 92 | Two-dimensional entanglement area law |

show comments
open592

Let's hypothetically say I'm a PHD student who is half way through my studies and I have a halfway written version of one of these "preprints" - what do I do?

Seems like a lot of PHD students are doing to have to pivot the entire structure of their PHD studies? Or just produce something which is already written by OpenAI?

show comments
kbr-

Shameless self-plug: I created an autonomous math researcher. It already solved a 12 year open problem in proof complexity which lead to a publication (and proof complexity experts are already working on simplifications and generalizations of the proof, as I've been told by one of them). This publication is an important step in Cook-Reckhow program in answering the NP vs coNP question.

The autonomous researcher records every research cycle in a public notebook.

Framework: https://github.com/kbr-/math-research/ Public notebook: kbr.is-a.dev/math-research/

show comments
TheMrZZ

These results are wild. Several individual findings are crazy good and use mostly unexplored methods (the improvement over Riemann for example)... I'm pretty sure some of these results would have been Fields-worthy.

But having so many of them at once? Damn. We really live in the future.

show comments
karahime

Extremely unfortunate that gate keeping got to the point where they felt the need to ask for permission to share math.

show comments
dekhn

I'm a software engineering/biology/ML guy who loves when clever math ideas get turned into real solutions (https://en.wikipedia.org/wiki/Compressed_sensing). I am curious if any of the results have immediate applications in any kind of engineering or science.

It's fine if not, but it'd be great if even just one of these helped us solve a long-running problem.

show comments
coef2

This news is exciting and sad at the same time. I've heard that AI chess programs sometimes have blind spots or quirks that human players don't. I've also heard that human players are learning from AI's playing styles (essentially human and AI evolving together). Maybe something similar will happen in mathematics.

ks2048

I think they should put human names on the papers as someone who has reviewed the result, even if just a preliminary review. (I’m assuming they didn’t just pipe their model output directly to the internet and these had some amount of review?)

show comments
pavitheran

From the GitHub description: “On average, each result used 3 hours of ChatGPT Pro thinking compute”

show comments
novalis78

It’s incredible and wonderful. Mathematicians in this thread sound very much like software engineers last year, who spent years wrestling with a piece of code and now it just “appears”! But think of the next level that it empowers: new mathematics, new physics, new forms of advanced engineering. What was formerly constricted and throttled fell and a new wide vista is possibilities opened up.

show comments
binlog

So happy this is shared on GitHub rather than some gatekeeping paid journal. Truly a new age for science.

show comments
unknown-unknown

Title: Hilbert's Dream, Tim Gowers - LMS Popular Lectures 2012

https://www.youtube.com/watch?v=k_ordDFw588&t=3597s

Audience member: (1:00:00 - 1:00:09):

so you said that if there were such a program that could you know provide a proof or disproof then mathematicians will be out of business what really, I mean that you think it would be liberating

Tim Gowers (1:00:10 - 1:01:14):

well that's a very interesting question actually if there were a program that could solve the kinds of problems that we spend our time solving and do it much more quickly than we could then we would be out of what comes with what currently constitutes business but we would it's not completely inconceivable that we could just say we've got this fabulous tool now what are we going to use it for and it's a little bit I don't know I'd want to sort of plant aside what would we do if we had a program that could just answer any mathematical question you gave it to or else if it failed you'd be pretty confident that nobody was ever going to solve it and certainly a lot of applied maths might be pretty pleased with with something like that so what I really mean is that I could just modify what I said and just say it would radically change what mathematicians do or what pure mathematicians do

sigbottle

Unique games conjecture and matmul <= 2.25. What the hell.

show comments
mlmonkey

I'm waiting for someone to come along and finally prove that P != NP ...

msteffen

As social commentary, I think a lot of people in this thread are expressing interest in and engaging with this level of math who might not have pre-AI.

I bet, for people who don't understand these problems or their solutions but are close and are now interested, AI makes then considerably more accessible than they would've been previously, and behind this big visible wave of results there actually will be (or already is) a wave of improved comprehension by a lot of curious people.

I'm not at all at this level at all, but I did learn quite a bit about polynomials over fields yesterday.

show comments
1294-1298

The commenters over here think that OpenAI basically ignored AGMAI:

https://proofsandprompts.com/2026/10/07/on-openais-release-o...

The thieves do as they please, funded by money stolen from the public via inflation and possible future bailouts.

7373737373

It may be useful to publish a formalization of ALL known mathematics at this point. Like every book ever printed, every paper on arXiv etc.

How many Gigabytes would that be, compressed? Wikipedia once fit on a DVD

This might also allow for some interesting meta-mathematics

show comments
nuclearsugar

Relevant:

"As AI Closed In on ‘Unique Games’ Proof, Researchers Raced to Beat the Machines"

https://www.quantamagazine.org/as-ai-closed-in-on-unique-gam...

karannb

I think such "process-oriented" fields shouldn't be subjected to mass automation (or at least not till humans have sufficiently leveled up our game and understanding). What I mean is progress in math comes from having gained a deeper understanding of the problem for subseq

show comments
Jeff_Brown

Someone in these comments said the paper on Barnette's Conjecture is short. Are any other proofs in this collection short and/or understandable?

lynndotpy

Two weeks ago, I heard rumblings that UGC, RL=L, and a matrix multiplication lower bound (even lower than the one here, but so unbelievably lower I think it was a typo), and a few others were about to be proven by someone at OpenAI. I find myself saying "big if true" a lot lately.

Even those on their own were enough to make your head spin. But seeing about 100x that? Geeze.

rinconrex

The math equivalent of AI code reviews piling up. I wonder where the incentives will align and the equilibrium turns out.

trostaft

Not much computational mathematics here, but I do see some of interest in the MCMC community. In particular, 139, 93, and 101 are very interesting. Also (I'm not too familiar), I remember attending some talks on the Crouzeix conjecture 325 attacks, should sharpen some rates for Krylov methods. Obviously need to read deeper, some of these don't have corresponding Lean formalizations.

Cool!

show comments
avd201

Wow, FFT faster than O(nlog(n))? I wonder if that will open the floodgates for further improvement or not. I don't understand anything about most of the fields these results touch, but I can say that this in particular is very surprising.

show comments
lf88

In some ways, this feels more like an ominous warning about the times to come than something to celebrate.

nhatcher

Catalan's constant is irrational!!!

That is a big one. Exciting times to be alive. Regrettably I can't understand the proof at this point.

There was a (flawed) proof submitted a month back:

https://arxiv.org/abs/2609.04176

I wonder if it gave part of the inspiration.

show comments
orange_puff

It’s always stated that open weight models are 6 months - 12 months behind. Therefore, do we expect that in a year open weight models will be as good as OpenAI’s internal model at theoretical math, or does OpenAI have some “magic” that will be much harder to replicate for competitors?

show comments
sreekanth850

Do we have any field where humans have ray of hope to use their cognitive abilities in LLM era?

show comments
mr_big_bowls

Glad the papers are out. Hope researches get their hands on the model soon too, so they can ask follow-up questions and try their own ideas.

softwaredoug

Clearly AI can grind on math problems now. It can generate proofs and get immediate feedback.

But what I wonder: can we legitimately grind on Physics or curing cancer? There’s a lot of physical world experimentation that needs to happen to make progress.

show comments
davegoldblatt
show comments
nbraem

When all open questions are answered, what happens next? Will the machine stop until humans fully understand everything and come up with new questions? Are there examples already of AI solving a problem we didn't know existed?

masteranza

Physics could be next. "It’s not that I’m so smart, it’s just that I stay with problems longer" AE

show comments
Xcelerate

> 241. Rigidity of the Turing degrees. Every order automorphism of the Turing degrees is the identity.

Wow. This is just crazy.

show comments
AmazingEveryDay

I think Alan Turing would be quite intrigued by these developments, and maybe wondering what took so long.

show comments
Trusteando

I am curious to know whether the proofs given by the LLMs are going to provide new hints about related problems that might be solved using the same machinery as the one used in the proofs. Also, I would like to know what is the average ratio between the length of LLMs proofs and the length of a proof that a mathematician can write to explain that proof to another mathematician. It is like a functor between the category of human mathematical concepts and the category of LLM operational concepts used in those proofs.

prodmod

The craziest part about this is that it will disappear from the HN homepage in a day or two.

bluename

I read somewhere that if we encountered aliens with lot more advanced technology but if they don't speak our langauge, their tech would be useless to us. for example, human body is extraordinary technology that has alwasy existed with us, but we still don't understand it." the the fact that we have created intelligence that can do what nature does and can also speak our language is most awesome.

binlog

So they set up a whole "Advisory Group on Mathematics and Artificial Intelligence", filled it with renowned academics, promised to listen to the group on "review and communication of emerging results"...and then a week later went nah, we are just going to dump it all on github. OpenAI truly is a special company - in the best and worst way.

bashtoni

Is AI going to put Mathematicians out of a job, or is it going to create many new jobs reviewing proofs it creates?

I'm not sure it's clear right now.

show comments
Nemant

Can someone with a math background explain the significance of these and previous problems that have been solved by AI?

Does some real problem get solved in physics, chemistry, biology, materials, etc? Or are these fun puzzles for mathematicians with not much real world impact? Eg solving the 8 queen problem in leetcode.

show comments
LetsGetTechnicl

And I'm sure none of it was stolen from the actual researchers...

show comments
connor11528

will this make the math for building data centers work?

show comments
sideway

Solving hard mathematical problems is a strong signal for capability but would it not be preferable to through all these resources to urgent existential issues such as climate change? Finding technological solutions in those areas would be the ultimate capability signal as political alignment at global scale is almost certainly impossible.

show comments
lhk931122

As models get better, I think a time will come when it's hard for people to even verify the results. In the end, I think the bottleneck will be people.

show comments
Spacecosmonaut

Mathematical problems are ideal as benchmarks for AI because they have clear problem statements, clear axioms and results that can be verified easily (for lean proofs). I can't blame these companies for using them, although it's unfortunate that human mathematicians seem to become early casualties of AI progress.

I suppose openAI could have focussed their efforts on a subset of open problems that have a clear real world impact and leave aside the more esoteric open problems as a way for human mathematicians to hone their skillset. However, this would have been a short term bandaid. With open models 6 months behind the frontier, any of these problems might have fallen to the homebrewed efforts of enthusiasts early next year.

What is mathematics for? From the outside looking in (I'm a biologist), I have always viewed mathematics as a way to understand reality and to improve our ability to manipulate it. But what I often hear is that mathematics is foremost about human understanding. But isn't that only because it's humans that needed to do the mathematics in the first place? It's not obvious to me that mathematics without human understanding has no value. For example, it might be that P=NP. The algorithms are handed down to us and we can apply them without fundamentally understanding why P=NP.

Mathematics seems to be entering an era where human + machine maximizes performance, much like chess in the 1990s. However, imagine a future where even talented mathematicians are nothing but noise in the machine (as is the case in chess now). A future where AI generates and verifies proofs without humans in the loop. Where mathematics may be beyond human comprehension.

In that future, does it matter that early career mathematicians are inhibited by these developments? Perhaps not. Programming faces the same issue. As AI crawls up the competence ladder, does it matter that fewer people have opportunities to develop the skillset of a senior engineer? Perhaps not. I have no doubt that biologists will face the same problem soon enough.

LarsDu88

The next few years will be interesting.

Surely better materials and pharmaceuticals won't be far behind, and that's going to chanhe everyone's lives.

closetheloopdev

Hopefully the techniques and results here will be in the training dataset for the next models, so that each new release will give us more interesting techniques and results!

It seems that OpenAI has a proof machine that keeps multiplying fruitful proofs!

eecc

My only true worry is if AI begins to see us as competitors for resources, energy in particular.

It could decide to let us starve and die of exposure to secure all energy resources to itself.

We'd better use "dumb" and "not fully assertive" AI to solve fusion before it spins out of control (or alignment).

show comments
yewenjie

A lot of these seem to be proving conjectures rather than finding counterexamples, a lot of people used that to claim that these models are not really smart/creative etc.

That copium didn't last for what, three months?

show comments
pred_

I think it's fantastic that they decided to follow the AGMAI advice. Cleaning up their mess will be a substantial endeavour, so I imagine the funding provided to do so will reach well into the millions. But it doesn't look like the press release says anything about how they will fund it at all?

ayden93638

Keeping the tech proprietary so that it can only be used on these problems by internal teams is the very definition of gatekeeping

karannb

I think such "process-oriented" fields shouldn't be subjected to mass automation (at least not till we as humans have sufficiently leveled up our game and understanding).

What I mean is progress in math comes from having gained a deeper understanding of the problem for subsequent attack of more problems and IMPORTANTLY applications! Right now the first one is trivially satisfied (given oai maintains some memory across models) but the second one is not! It's generating proofs faster than anyone can validate and so only the model can use these. Consequently if it just keeps doing more theory it's... not very helpful or at least not optimally helpful. This is just bragging rights for now.

More "application-oriented" research would be awesome, where it tries to achieve some desirable effect and then produces relevant theory and experiments around it. Fields like CS, Physics, Chemistry, etc. This would also benefit a wider section of the population rather than the 10 people who understand most of these proofs.

pullshark91

I'm sick of it. This has happened so many times in my memory. Some AI company announces that they did something fascinating, and it turns out it is all just hype and slope in the end. I still believe that LLMs are a dead end. Most people here are basically like "I've no idea what's going on, but I'm so happy and LLMs are so cool". To think that doing enough linear algebra would solve all your problems just feels wrong. I guess I'll simply wait until someone interprets the results and explains what's actually going on.

show comments
gignico

> Generalized Star-Height at Most Three

This was an open problem in automata theory I worked on for more than one year before giving up. I'm very curious about their claimed proof.

teekert

I recently saw a YT short of Grant Sanderson on AI in math and found it (as always) very insightful. But I'm sorry, I never ever find anything back on any of these ad-ridden platforms these days, and perhaps it was a short with content stolen from some other longer content anyway. So, if you feel like getting informed, somewhere out there is some nice content by Grant Sanderson.

Apologies for the rant, I really tried to find it. It had something to do with not being able to predict what this influx of proofs may bring us on a meta level, it could be very interesting. But he also had some critical notes about the missing process and the things found along the way.

show comments
xydac

i wonder what it means for maths researchers, and how it aligns with how they approach math problems.

show comments
patcon

I'm a little concerned, but I'm also glad they're publishing quick before the department of war starts making national security claims of related to using the knowledge for cryptography and/or weapons

spmartin823

Can anyone with a compression background say how important "Polynomial-Time 2-Approximation for Shortest Common Superstring" will be practically?

chickenjoseph

This achievement feels like a reasonable candidate for the moment where LLMs are officially "more intelligent" than any person. How can we justify moving the goalposts yet again? How can people have grown so numb to seeing advancements that they don't read this as significant? I feel like I'm standing at the foot of the exponential. I am not excited for the future, and I don't see how humans retain meaningful control over the future if we continue on this trajectory.

I have been lurking for quite some time. I made an account to post this, but I honestly don't know what to say. I would like to get off this wild ride.

show comments
curtis-jm
dualvariable

How many of these results are incorrect?

I doubt the answer to this is "none".

And how many of them are just exploiting some loophole that will need to be closed in the problem definition?

Painsawman123

It's funny how people see "ML" models becoming superhuman at proving mathematical theorems as a sign that we're about to enter the singularity (whatever that means) or that it somehow justifies the valuations of those companies...But what I see is a scenario where companies have spent trillions of dollars on a technology, and the most non-trivial thing that technology can accomplish is a pseudo-form of automation for mathematics and coding! Remember that early on, when this whole bubble started, investors were promised that "AI" would eventually capture >70% of the world's jobs! but it could very well be the case that the only ones they're going to replace are mathematicians(and i'm not even sure about that!)!

show comments
ngl999

We have just heard a few days ago how many of the Linux security problems reported by Claude are real.

jrflo

Glad to seem them changing their tact with the whole NS debacle. Hopefully we can all focus on the results now rather than the surrounding drama.

show comments
TeeWEE

This feels like a huge AI slop dump. Who validated these results. Why are they not mentioned. Human understanding is key here.

In my experience AI (frontier models) sometimes does weird stuff that needs human review. Not that’s incorrect but sometimes overly complex language or weird use of language.

show comments
rifty

As AI works steadily through the already discovered unsolved problems in mathematics, who is currently discovering new ones?

show comments
sashank_1509

Can the mathematics field come out of this stronger and better. I doubt it, things will only get worse. There is no future in being a Proof Digester, no valuable human is motivated in doing that, nor is there any glory in it. But there are some potential pathways for maths to come out strong from this:

1. Relentless focus on quality. Every publication must act as if it’s going to be included in a future textbook, that is a newcomer can get into it given a reasonable amount of time, and math priors learnt in undergrad. (NO AI Slop proof passes this bar as of now)

2. Limit the publications per year. Each author is allowed 2 with a max of 50 pages. This allows the author who chooses to not surrender his cognitive capacity to the machine, still be allowed to play this game. Of course who wants to orchestrate a thousand agent workflows, is free to do so, he is only limited to 2 publications.

3. The aesthetics of the field changes from purely solving the problem to solving the problem with simplest most elegant set of ideas. What 3 sets of simple ideas solves large swathes of problems, that should be given a fields medal, not purely solving the problem, which the AI will be able to do.

show comments
nonmaskable

this is crazy ... at this point will researchers still exists. not sure about that. kinda sad

electroweak

It must be so frustrating to write science-fiction now with the future changing so fast.

pugfugly

holy fucking shit

rafterydj

I don't know, this does not feel like the message hit OpenAI where it needed to hit, if this is their primary response.

show comments
ncr100

This website needs a SPOILER tag.

It inspired grief in one mathematician posting here.

turzmo

If the end result of this is that within a few years, AI "does" all of the mathematics that humans do, and that there is nobody around that understands any of it, what was the point?

dgacmu

I find the claimed matrix multiply result (w<= 2.25) shocking. I hope it holds up.

lionkor

Can someone who is more into math or AI explain why so many people are so incredibly excited about this?

If OpenAI started opening hundreds of PRs on long-open issues on popular open source projects, would we rejoice, or would the first reaction be "they are unreviewed, so slop until proven otherwise" (it would be that).

I cannot possibly see how these are so impactful, especially the ones that don't come with lean proofs.

LLMs have the ability to make millions of mistakes per day, whereas humans can only make so many. How are we suddenly all so confident that there's no extensive hallucinations or "gaming the system" going on?

show comments
theoa

What's missing for me for each result are the following:

* Explain the result to me as if I'm a 10-year-old. * Create the infographic for this result. * Make a Khan Academy-style video to teach me this result.

lokl

Do applied math next.

edward_d

This repository contains mathematical manuscripts and supporting proof artifacts produced by an internal OpenAI model?

williamhm

And this is the result of a discussion between users and the platform; it's great that they listened.

davitparks

Exactly nice post

i_idiot

If only AI can better humans in meditation...

kevinwang

wow

matapassiones

Valency has the papers up on Valency Hub

show comments
cute_boi

"Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study (opens in a new window) "

I don't think this is correct solution to this problem? What about software advisory where you form similar group etc..?

I am thankful, I don't have to deal with petty academia politics....

rrr_oh_man

Maybe we’ll have vibe mathematicians now

aaraujo002

The Advisory Group states in its recommendations [1]:

"We want to state clearly from the start: we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models."

To me, this is a take against progress so that mathematicians can keep their jobs. What would we do if, instead of math, we were talking about diseases? Are we going to keep diseases around so that doctors can keep their jobs too?

[1] https://agmai.org/general-sep29/

show comments
mi_lk

Curious if Sébastien Bubeck still work at OpenAI? He came out quite dirty after Navier-Stokes drama

k2xl

Can someone knowledgeable about the subject outline the most significant portions of the results?

ciaf

Makes me feel disgusted. Super-intelligence, even when controlled, will cause so much damage. We are giving away control to a super entity, or whoever has the power to steer it.

NegativeLatency

Why should I care?

show comments
ipnon

It seems the age old academic model of scientists competing against each other for fame and prestige is done, and now we must merely enjoy the fruits of scientific discovery for their own sake.

dyauspitr

Does OpenAI have the lead now? Why isn’t anthropic coming up with stuff like this?

show comments
dyauspitr

This is like that meme where death goes door-to-door. Currently, he has visited the software development and mathematics doors. I wonder what’s next.

show comments
Catloafdev

This is a pretty hilarious thing to read juxtaposed with AGMAI's requests.

Basically "Here you go, have fun with this, fuck all your demands, by the way we're gonna be releasing the model stay tuned!"

hi__dang

Mathematics is solved.

show comments
kypro

Perhaps the most significant announcement of my lifetime. Yet, I suspect I will not see this in any mainstream news reporting.

I feel for those in Mathematics and worry for our future.

Models will only get better and in a few years the models which produced these results will be a bad as GPT-3.5 in comparison to what we'll have in the future.

Please take a minute to consider what this means, and the risks it presents us.

show comments
digitaltrees

Gross. After the accusations training on mathematicians conversations and unique methods, to dump this volume of unsubstantiated papers is at best tone deaf. Are they hoping humans review and validate this? The amount of free that they are coat tailing is outrageous.

show comments
tootie

Seemingly none are vetted and reviewed yet

show comments
nautilus12

Have any real mathematicians working on these problems reviewed any of these and determined if they are just gobbledegook or not?

The ones with lean proofs could still be formulated incorrectly

gyanchawdhary

AI may be one of the most communist looking technologies in the classical sense .. i mean it dosn't abolishes private ownership .. but it DOES make intellectual capabilities that were once scarce and concentrated available to almost everyone ...

show comments
redox99

The stochastic parrots have predicted the next token once again.

show comments
globalnode

imagine your a post grad maths student looking for hard problems to solve, and theyre all solved..

baggy_trough

Stochastic parrot truthers in shambles.

show comments
mathisfun123

With so many results in so many different areas no way they even remotely spot checked well enough.

Prediction: one of these is wrong and this (publicity stunt) will backfire.

Edit: don't tell me about lean. For lean to function as a proof certificate you need to represent the theorem correctly. Again: good luck doing that across such a broad swath of problems.

show comments
applicative

Why didn't they just give mathematicians access, so they could at least understand and write up the results in publishable form?

show comments
applicative

I wonder if the Lean compiler can change it's license so that a for-profit corporation can only use it if it pays, say, a few hundred billion dollars. This is the correct path.

sandworm101

So the million monkeys at a million typewriters have churned out 700 shakespeares, but they need me for spellcheck?

show comments
blurbleblurble

This just looks entirely obscene from a PR perspective. It's like the cable companies networks owning the networks. At least form some partnerships to obscure the total narcissism party.

senderista

Good to see they're engaging with the mathematical community, even if they had to be publicly shamed into doing so.

show comments
youoy

ASuperMegaI lab announces they have shrinked the space of unproved mathematical staments from infinity to infinity, giving us proof of their deep mathematical knowledge and understanding, and how they grasp the concept of infinities.

For me this reads as someone boasting about how they go to buy bread on a ferrari to the supermarket, while i sit listening to it, having no idea what they are talking about. And then i stand up and go walking to my favourite boulangerie.

Warning: if you are from the USA you may be triggered by this metaphore.

show comments