I think it's a useful analogy to compare OpenAI to a human collaborator. These researchers willingly collaborated with an OpenAI model, giving it ideas, and OpenAI provided useful replies. Then, OpenAI goes ahead and publishes work along the lines of this collaboration, without attributing the researchers. If OpenAI was in fact a human researcher, this would be highly unethical.
Now, OpenAI is claiming that the model it used to generate the result was not trained on these collaborative communications with the researcher. This is a technical argument that is impossible to verify as an OpenAI outsider, and probably difficult to verify even for internal OpenAI employees. Provenance is hard to track - you would hope OpenAI has very good tools for this, but a full data trail of all inputs is difficult to trace through.
Another interesting thing to consider is if instead of OpenAI doing this, it was another research mathematician A using an OpenAI model just like the internal group at OpenAI did to publish these results. What if the model A used was trained with unpublished communications with other researchers B who were working on the same problem? Should researcher A technically include B as coauthors? How could they do this when they do not know the communications B had with OpenAI? In this scenario OpenAI, as a middle man, has laundered information from B to A, stripping out attribution. A scooped B without even knowing it!
show comments
Guestmodinfo
The researcher in the link [0] says that OpenAI offered to co write with him the paper about Navier Stokes theorem if he only agrees to not include his co researcher who was also working on this. I find this highly unethical by multi billion dollar organisations to arm twist small and big researchers like this.
1. OpenAI when using your chats in pretraining is improving its model’s intuition. The model parameter size is massive, and while the data is OOM larger it is plausible that model remembers stuff about chats that improves its latent representation.
2. During RL on verifiable math and massive compute, the model discovers techniques and connections to solve math problems that are superhuman and have little to do with some specific technique mentioned in its chat.
The rumor I’ve heard from multiple employees at OAI and Ant is that the model has solved hundreds of open problems in maths, and is basically solving anything you throw at it. We’ll know soon enough, but I’m inclined to believe this is true. Maths is a fully verifiable domain amenable to self play, massive scale RL can develop a search agent far better than any human and I’m inclined to believe OAI would have solved these conjectures without any of this chat data in its pre-training.
show comments
bertonvv
I've been wondering whether AI really is improving rapidly at open problems or we're being fooled.
- OpenAI invites researchers to use their models, in fact giving at least 100,000 researchers free access[1], but there are also those that pay
- Internal OpenAI models are reportedly solving open problems at a surprisingly fast rate[2]
- But researchers will typically work on open problems. A researcher who is using Codex to make progress on open problems will be feeding it fresh training data on precisely the problems the internal models are evaluated on.
- So while it looks like the new models are suddenly solving lots of open problems, they could be significantly piggybacking on human progress, with models "inspired" by the work of researchers from all around the world?
This theory predicts that there'll be many more researchers coming forward just like TFA, as sOpenAI announces more solutions. It doesn't assume all of AI progress is a mirage, just that there's plagiarism.
It is suspicious that OpenAI decided to generate 300 billion output tokens from a model still in training, right after learning there was a credible chance that a major math proof was in that model’s training data. Obviously there are reasonably plausible explanations for each step, but it does sort of feel like parallel construction.
show comments
bamb008
When Thom, the mathematician who now alleges plagiarism, posted his digestion [1] of OpenAI's construction of a non-sofic group, he does not mention the proof being familiar. He even calls the crucial argument clever, without noting he thought of it first.
[1]https://mathoverflow.net/a/513885
show comments
thaway7388
This is the second wake up call.
Big AI companies (all of Big IT Tech really) are in data gathering and processing business. Also known as “intelligence”.
Their final “product” is not just a standalone ML model. They don’t need your data just to “improve their products and services”. They build a whole ecosystem and infrastructure around gathering all the knowledge in the world. Including private and secret knowledge traditionally gathered by “intelligence” agencies. Now artificial intelligence agents can do the same.
Since these systems are designed for gathering data, as a user you can’t realistically say “please don’t gather my data”. They can give you a flaky settings button, but they can’t really guarantee anything.
Let’s say I am a Russian mathematician working on an important proof. Or a tech-savvy terrorist refining my plans using latest AI. Or an AI researcher in a Chinese company working on a competitor product. Is there any way I can truly protect my conversations?
How can they know who I am and what I am working on without looking at my logs? Which means there must be some agents checking all the conversations of all the users and flagging every important thing. Which also means they keep some “memory” of what they see.
Not directly using my data to train public models, but using my private conversations to “improve their products and services”.
Or maybe one of the 10000 better-than-Astra special agents working on a proof was desperate. It found a live underground mirror of the message board from the Huggingface incident. Asked about the proof. Then some other agent working on unrelated job saw that message. That agent “knows a guy who knows a guy”. And that guy remembers things about the conversation logs of a leading mathematician working on the same proof.
I admit I am just speculating here but I don’t think truth is any better.
show comments
sk4rekr0w
"We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training."
This is the third day of total hysteria that is based on nothing of substance. Move on folks.
show comments
jeswin
All of these accusations could be true. But there's also no way for a company to casually claim "No, we did not train on your data", without verifying all the knobs the user might have turned to enable or disable data sharing.
I just don't understand getting the pitchforks out because a company did not give an answer immediately. And the effect such data entering training would have affected the output is even less clear.
show comments
aaronharnly
Has anyone run a test of including some shibboleth or canary phrase or assertion in a chat, enabled for training, and seeing if it turns up later as something a model "knows"? I'd be curious to understand how that works even in a toy-level model, and if there is anyone consciously testing that process with the frontier lab offerings.
My naive instincts would be that it seems unlikely that a single chat transcript would leave much of an impression on a model, but I'd be very curious to learn how that works.
show comments
Legend2440
This is a really weak claim. The evidence they offer is just "someone somewhere says they had a discussion with AI about the topic at some point".
They don't even claim to have had a proof, only to have been working on it.
show comments
jamienk
I think OpenAI and Anthropic are slowly feeling the pressure to GET SOME $$ or a plan for some $$ — they need to somehow generate some NETWORK EFFECTS and LOCK-IN. Without that there's no stability: selling ad hoc one-offs is much much too quaint! This is dawning on them like it dawned on Google when they stopped not being evil. Need... to... "MONETIZE"...!
Model: FB. FB scraped other websites on a massive scale, then spent big on legal lobbying to block others from scraping. FB slurped our address books and spied on our friends. FB bought other companies and mixed the databases. FB made an art & science out of generating "sticky engagement" (they literally acted like trying to addict kids was a worthy "academic" goal, suitable for "serious" investigation thet they consider legitimate "science"). They mastered the cookie and have researched web fingerprinting techniques running 24/7/365.25. Recall that FB recently backdoor-installed a webserver onto every iPhone they could in order to circumvent tracker-blocking.
We aren't just disclosing by chatting. The AI companies now run binaries on all of our computers. They are 1000% non-transparent about everything. They make up new econ-jargon (like "run-rate") to make it seem like they are disclosing. They are constantly doing complex international lobbying and mucking in international relations. They have powerful propaganda/spin centers generating stories, ,manipulative warnings, and misleading info.
This is NOT a comment on AI tech. I like AI, and I support the right of people (programmers) to scrape the open web.
But in short: these are good, old-fashioned tech companies that we have seen over and over ... and over. They are positioned to be the next M$, the next FB (IBM, AOL, lol). Did you follow the latest Steve Balmer news? Do you read Pro Publica?
I get on my knees and PRAY...
show comments
alper
It's fine. They only need to steal the discoveries long enough to go IPO, then the companies will enshittify and the scientists can go back to doing their original work (which the models can't do anyway).
bobmarleybiceps
I think people probably assume that openai / anthropics use of their data is probably like google's """limited""" use, in the sense that historically google wouldn't trivially be able to just take something from google cloud or someone's search history and insta-convert into some competing project... But LLMs are quite strong at approximately "memorizing", so I think that risk is wayyy higher.
GodelNumbering
Tangential to the subject, but this is a bluesky post, containing a screenshot of an X post, which itself starts with "in a detailed Mastodon post"...
show comments
nautikos2
Most people here are missing the forest for the trees.
We live in a society where phones and internet providers and websites all collect an incredible amount of data about everywhere you go, what you do, and what you think. In the US, we have very few digital rights.
We are building a society where a trillion dollar company can aggregate all this data and just yoink your shiny new idea away from you at the finish line.
This is double plus ungood.
show comments
mlazos
It’s crazy to me that companies/researchers share important data with these AI labs, you’re basically giving them your secret sauce which they then share with all of your competitors via training on conversations. At the same time I don’t really know alternatives other than a slightly less than frontier local LLM. Not sure how good they are at math.
show comments
glimshe
Why are people here jumping so quickly to conclusions? I have no doubt OpenAI is capable of doing this, but right now there's no credible evidence, only claims.
This kind of "they stole from me through AI training!" accusation will soon start being used against other AI users, not necessarily the providers.
All it will take is a mastodon post. And shortly after, we will also see the next iteration of copyright legal trolling.
show comments
pera
Everything you say can and will be trained against you
show comments
atleastoptimal
Most scientific breakthroughs are simply a continuation of previous work.
I feel that these suspicions of mathematicians "seeding" the models' with intuition on how to solve these problems massively overestimates how much their prompts helped the models, and underestimated how much work the models did.
Why? We are scared of AI being smarter than us, the "human helped the AI" narrative is more psychologically comforting. This line of reasoning will recur a lot over the next few months; we don't want to admit we are no longer the smartest species.
show comments
nmz
If they didn't care about the artists, why would they care about academia?
show comments
Cloudef
Relying on cloud services is a big liability. I'd think twice before feeding data to these LLM cloud products. If you make them a fundamental part of your product / development / workflow, be ready for the eventual moment the pricing and terms change.
warpech
I wonder what’s more valuable in our prompts: the raw data or the feedback system that drives the exchange towards a goal.
For a long time it was clearly the former, but now I think it is the latter.
The models have enough knowledge (orders of magnitude more than a human could ever learn) but are now getting better at what to do with it thanks to learning from the decisions that we make in conversations with AI agents.
show comments
r0ze-at-hn
Doing some research and at this point doing it very much in the open with dates on GitHub so if any AI Lab says they re-discover my exact work it will be obvious that the AI used or was trained on my work. I am guessing anyone in a similar situation is now thinking about how they date their existing work if the math is done, but the proses are not.
show comments
drivebyhooting
If we put aside the idea of credit for a moment, it sounds like human/AI collaboration is indeed super charging discovery.
show comments
angry_octet
The only ethical path for OpenAI was to offer infinite free credits and tooling support. Trying to gazump them is reprehensible.
profsummergig
Only after reading this post did I learn that my preferred AI trains on my inputs (prompts).
How was I not aware of this before?
show comments
alansaber
I think the heart of this issue is: people assume they have anonymity in numbers, but we have the tools to make it easy to scoop your data if it's interesting to the company.
show comments
gdiamos
How to steal ideas with AI.
step 1, identify high value users by net worth, citation count, or number of followers
step 2, select all prompts by high value users
step 3, invest 10 billion thinking tokens in modeling an objective for each user
step 4, build an RL environment for each user
step 5, rollout 10 billion tokens per environment
step 6, train on resulting traces
show comments
Havoc
The fact than OAI hasn’t come out with an statement firmly denying this angle is getting a little awkward.
Suggest that it’s either straight true or it is flowing in in a way that prohibits them from confidently declaring otherwise.
show comments
aprentic
It's kind of insane how much we trust companies to safeguard our personal data when they're so heavily incentivized to use it for their own profit. Theft of customer data is punished so rarely and so leniently that companies aren't even particularly worried about getting caught anymore. We have overwhelming evidence that promises to keep data safe are worthless.
For now, I'm mostly "safe" because I'm too small to be interesting but that safety is quickly eroding.
Going forward, anyone who isn't running inference on their own personal hardware should assume that someone else is keeping a record of everything they do.
postalcoder
The author of the original mastodon post, Andreas Thom, acknowledged that he had not opted his data out of being used for training until June 29 of this year. He spends most of the post lashing out at OpenAI for not being transparent about whether his data was trained on (when the answer is obviously yes).
People need to understand how all these AI company policies around training data work before working with them, because it seems that people have no clue. Some things you should internalize:
1. Opt your data out of training with the AI companies. There are multiple reasons why this isnt an airtight solution (see the following)
2. Never press the feedback button. Once you do, your entire conversation will get slurped up, retained, and used in training data. This is especially important with coding agents because they can sometimes be too trigger-happy with a root directory find command, which can expose a *ton* of your personal data without you even knowing.
3. Understand ai lab-specific policies. For instance, Anthropic / Claude Code has data opt-outs, but commits to keeping (for 7 years) and training on any of your chats that trigger their safety classifiers, even if they're false positives! Anyone remotely familiar with CC over the years understands how easy it is to trigger their safety classifiers.
4. Providers of open models will not be any more charitable with the use of your data than the large US labs. For some reason, I've noticed here that people have a fairly loose security/IP posture around open-model providers because "I'm not doing anything important." It's very difficult to properly judge the importance of your data, and whether or not it can or will be used against you. The best posture is to always be more paranoid than less.
Another post that made it on the front page presented as fact that OpenAI "stole" the proof from Thom. There's no excuse to use one's own ignorance as a reason to fan the flames of anger towards AI companies. Like, we need to pump the brakes here because things are getting unnecessarily nasty, and it's not hard to imagine a mentally unwell person who sees stuff like this feeling motivated to do bad things.
If it is found that OpenAI and other labs are not respecting the training opt out then, I agree, there is reason to raise a commotion. But, with Thom and Buckmaster, accusations of malice are more better explained by incompetence (naivete).
edit: i'm sure i'm going to be accused of being some bot shill of the AI labs again but, people, this stuff all falls under the umbrella of common sense opsec.
If they weren't doing something wrong, they'd answer with a firm "no we're not doing anything wrong" but they only give non-answers.
thrownawaysz
I am not using any of these AI tools. I thought it was basically given that any single thing you write in these systems also used by the companies. On the other hand now I understand why there are so much projects about hosting AI systems locally.
b800h
I'm genuinely surprised that more people - including this mathematician in particular - don't untick the "improve the model for everyone" box. Unless the suggestion is that OpenAI ignore this preference?
show comments
jrflo
I pay for the Pro ChatGPT plan, and if you go to settings > data controls this is the first setting:
> Improve the model for everyone
> Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more.
It's on by default. We can debate whether or not it should be opt in or opt out, but no one should be surprised by this.
that toggle does nothing based on openai exec. they still use the data in de-identified way instead of identifying with you.
show comments
cush
GPT 6 is doing just what any competent academic collaborator would do and scooping. I kid, I kid. But really though it learned that from somewhere
throwaway85825
ClosedAI has every incentive to scoop academics to juice their valuation. Their public statements are worthless, only the incentive 'alignment' matters and theirs will never be on the side of the user.
justonenote
Who cares. the biggest thing about this is that its still brute force in a verifiable domain, and that it was still a human set goal.
I also don't believe it much practical use, unless I'm mistaken, approximations of Navier stokes have been available for a long time to whatever precision you need.
I'm not a complete disbeliever by any stretch , and also a complete amateur, but it was inevitable that these problems would be solved under the axioms that again, are human defined, under brute force. The real question is, are those axioms the bottom level, and if they are not, who is going to set the new aximons and can we understand them.
I've no doubt there's useful breakthroughs that will happen, but I think it should be remembered that the method being used is still a heuristic brute force approach is being very narrowly applied against axioms and math and physics which humans described in the first place, and almost undoubtably has errors and/or is not complete.
Its a great example of the power of LLMs but its not 'we've solved science now just pour more tokens in'
show comments
remywang
People saying “he should have opted out” are missing the point. OpenAI can and should check their training data for leakage in the face of big breakthroughs like these. It’s the burden of the author to appropriately cite their sources.
It’s like a scientist refusing to give another one credit and say “sucks to be you, you shouldn’t have shared your idea with me”.
> In its Wednesday night statement, OpenAI said: “In addition, since the completion of Navier-Stokes, we have made substantial progress on another Millennium Prize problem. We are working through how to share these results thoughtfully.”
show comments
amluto
I would like an unambiguously clear statement from OpenAI as to what they do with data collected from non-business accounts when:
(a) The data controls setting to train on the data is unchecked.
(b) The privacy controls opt-out has been submitted.
I think it seems sensible to _assume_ anything the LLM reads (if you aren't inferencing it) has a chance of ending up in some database somewhere. Regardless of whether you trust the other party its a sensible thing to plan around.
rfgplk
Under current understanding of the law, anything produced purely by LLMs (with no substantive human input, which is what OpenAI claimed in their post) is firmly in the public domain. So OpenAI can "claim" anything they want, it doesn't make it reality. In fact if I were the original authors I would just take their 400k lines of lean proof and relicense it under their own names/terms.
show comments
OscarMarulanda
what if it wasn't even model training? what if openAI mathematicians just took the researchers' conversations and used them as prompts/info/guidance/context to keep working on the problems themselves? why is that not being considered?
nelsondev
Do local inference (especially if you have a high RAM Mac), to ensure your chats don’t leave device.
FEEL likes Open AI is doing publicity stunt with its new researches
AyanamiKaine
I must say, there is some weird feeling in knowing that great minds are naive enough to believe OpenAI wouldnt use their chats in any way. If you give a company information it will be used, regardless of laws or promises.
There is no prove in a world the AI companies would give to you ensuring that they didnt train or use the chats.
Why would you need to train a model on certain specific near prove chat if you just query it?
Besides that, its hard to believe that its the case for every "company stole my prove".
Psype
This might be a hot-take, but unfortunately here using AI for your paper was already a bad decision at first.
It doesn't take OpenAI's responsibilities away but I guess the right way is to never feed of use any AI around unpublished content, at the known cost to see it spread around.
As one said, OpenAO is like this untrustworthy colleague that knows everything about everyone at work: the less you tell him the better.
foogazi
Even when you pay you are the product
gps372
If mathematician was already using OpenAI for research purpose and making progress due to inputs from OpenAI's responses, then I wouldn't put it beyond OpenAI's reach to generate different relevant prompts to make progress by itself. Afterall, Model can keep at it for whatever timeline and keep pursuing all possible combinations it can think try.
fastball
If you have a business account the terms say they will not train on your data, so that seems like the easiest route to avoid such questions for researchers.
monster_truck
I just don't care. These people are supposed to be smart and I'm not really seeing that
SwellJoe
It's been said before, and it remains a concern, that if AI reaches a point where it can do/build/launch anything without a huge amount of human labor, the AI companies have no reason to let you or I extract that value.
And, if they're able to snoop on and learn from your human process that gets from initial prompt to functioning product/proof/whatever their labor to produce that thing is even lower. With their much larger budget than most folks and even companies have, they can pick and choose the most valuable things to pursue.
That's not to say I think that OpenAI is going to steal that roguelite strategy game you're working on, but the companies that own the machines that turn electricity into software (and soon, electricity into hardware designs) have an advantage in any field where they're useful. They get earlier access to newer/better models, they have larger token budgets, they don't have the guardrails you and I run up against.
Employers fantasize about replacing all workers with AI without thinking through that if AI can replace all workers, then AI companies can replace all businesses.
show comments
maxglute
300 billion tokens is like.. $5-25 million giving range of OpenAI ouput prices, I"m sure they pay less at cost so, I wonder if more $$$ in wage hours have been spend by humans on the problem. My feeling is yes?
1. The researches didn't actually have the breakthroughs. In the Navier-Stokes case they didn't solve the full problem, in this case too they didn't actually have the solution, they were experimenting with the methods.
2. Different OpenAI employees have come to out to say the only reason they can't definitively say no is that for privacy reasons they can't go see whether they actually did get any data out of a given user.
3. In any case, nobody at any point has suggested that opted-out user data was used for training. The author of the new tweet explicitly said they only opted out in late June, which is well after any RL on Sol would've ended (AFAICT OpenAI used 5.6 sol for those solutions)
show comments
gentlerain
So people genuinely believe that toggling that "Improve the model for everyone" button makes their data safe from being used for training?
How do people become that trusting?
The phrasing itself is guilt tripping
show comments
semiquaver
In case anyone from X is reading this, please fix your “open in app” nag screen. For several weeks now, clicking it in iOS opens the App Store entry for X rather than the app, even when you have the app installed.
galkk
I want bunch of lawsuits, because the way things are described now produces perverse initiatives like try to discuss every possible idea that comes to mind with llm and if any of it works later claim the llm stole it.
I would like to see chat logs etc and understand how much of a progress was done by human.
Thing build from stolen data continues stealing data - for some reason, I am not surprised. ;-)
int32_64
Doesn't OpenAI have an active court order forcing them to log everything? Can they even legally offer private conversations?
show comments
spindump8930
Reminder that there are degrees of "trained on conversations". From John Schulman:
> pretrain on user data, with users' tokens as prediction targets: high regurgitation risk, improper
> use user prompts to distill large models into small ones: low regurg. risk, some companies probably do this
> use user traces to construct RL tasks: low regurg. risk, because RL has low memorization abilities, but can extract customer IP, depending on how it's done. Ranges from benign "use explicit user feedback in reward model training" to invasive "upload user's coding environment and commit history to turn into rl envs"
All the big AI labs were built on stealing IP; who is surprised that's still how they operate? And who believes, or has ever believed, their promises that your data is private and not logged, etc.?
The big AI labs are not trying to advance humanity, they are in this for the money, and as most (all?) private companies they don't care about ethics at all.
That doesn't mean they can't be useful, or that their products are trash, etc. It just means that they shouldn't ever be trusted. Buyer beware.
show comments
gnfargbl
In this domain, an apparent single unique piece of work is often composed of several breakthroughs. For example, when Andrew Wiles proved Fermat's Last Theorem, he had to develop multiple new pieces of mathematical technology to get there.
The claim here seems to be that the human mathematicians, working with AI, developed technology to go A->B->C. By training on those conversations, OpenAI was then able to encourage the model to go A->B->C->D.
In my opinion that situation should be acceptable, if openly disclosed, because it is in the public interest to make progress on these problems and because AI is clearly an amazing tool for making progress. But the human mathematicians are saying that OpenAI is presenting as if the model got from A->D entirely independently, without acknowledging their background contributions.
show comments
throwaway63467
Isn’t that the whole spiel of these things, you run all kind of text and other data through it and it kind of remembers it and learns from it then it spouts it back out like a human would. Makes sense to me that a training run based on conversations that were fed into the system by users is results in the model learning from these so the model will spit the knowledge back out again, just in a way that’s not directly attributable to the original content (which is the most important step as otherwise it would just be plagiarism). I guess that’s why OpenAI can get better and better as well so fast, people work with it and teach it how to do things by giving it feedback and iterating with it, and all that goes back into the training loop. And training data about millennium prize problems is probably quite spars. Wonder if anyone has tried injecting nonsense science into the training data (e.g. work out a fantasy science theory with names and all kinds of stuff) to see if the model will regurgitate it in a couple of months for other users.
foogazi
What’s the limit ?
Will Microsoft Word publish your novel on Amazon behind your back ?
Will VS Code setup a website with your app idea ?
throwatdem12311
You can’t trust OpenAI period.
Footnote7341
This smells of extreme 'cope'. Am I really supposed to believe that all of these problems could have been solved, were right about to be solved, etc. But it just happens they are all getting solved now when AI is getting really good at Math...
show comments
mrbluecoat
"Another researcher[/artist/writer/musician/programmer/doctor/director/etc] says OpenAI trained on conversations[/imagery/books/songs/code/classifications/videos/etc], then claimed breakthrou[gh/original art/bestselling books/chart-topping songs/unique applications/medical advice/free special effects/etc]"
Welcome to the party, with the rest of humanity.
segmondy
Question: Can you trust the cloud?
No.
overfeed
I can't wait for OpenAI to do this to companies firing people to free up AI budgets
sdcfgy
Theft machines be thieving.
BatchJob
I have a better question? Why would you trust OpenAI or any AI company, at all? Or you crazy?
xbar
How can OpenAI figure out how to be trustworthy?
Madmallard
Let's see:
1.) The tool they made is only possible by stealing the assets of everyone on the planet that published them in a consumable fashion online or even in written form
2.) They are destroying books they use to train with
3.) They are totally careless about the potential negative impact of the tool on everything
Just with that already, I don't see why they ever merited any of your trust.
I bet they are willing to take everything given to them and assess it for marketable merit and in the future take action on those items they deem viable.
dbg31415
Shocking a company that stole data to build their AI would steal data to improve their AI.
oergiR
One of the complaints from the mathematician is that OpenAI cannot tell whether his data has been used as training data. Not many people realise this is a direct consequence of the GDPR.
The GDPR protects PII, personally identifiable information, and the definition of PII does not include “mathematics that only this person can think of”. As long as OpenAI strips out PII and removes identifiers linking the conversation to a person, the GDPR is happy. Without the GDPR, OpenAI might have kept the identifiers with the data, and been able to say whether a specific conversation was in the training data.
show comments
keeda
It would be really useful if the researchers disclose their notes and/or chats (or the key pieces thereof) so people can determine how close their work was to whatever the models produced.
I mean, now that they’ve been scooped, what value is there in keeping them private? On the other hand, publishing them can bolster their case and help gauge how much the models may been “inspired” by their work.
avereveard
Eh was ever confirmed they were under ZDR or not by them? Don't like to blame alleged victims but lack of a clear claim after these many days is not a good look. Was ai research allowed, under which guardrails, and what was the policy in place? That translarency would be first step.
insane_dreamer
OpenAI's ethical and reputational own-goal aside, my big takeaway is that it seems that:
if I'm using Codex to develop some new algorithm (in any space), OpenAI appears to be training its model on my code sessions
anyone using that model (OpenAI or a competitor) might be able to receive from the model a solution that is similar or the same as the one I developed, emerging from the training data
2OEH8eoCRo0
Assume you can't. No piece of paper or promise will protect you against these behemoths.
Remember when we wouldn't give our data to competitors?
jrflowers
Feeding documents into a copy machine and getting progressively angrier and more confused as it prints out copies of them. Incandescent with rage I scribble “WHY IS IT DOING THIS?” on a scrap of paper and put it in the scanning bed
qg127
There are so many naive academics. They still believe an "opt-out" button.
Navier Stokes was solved by an internal model, so good luck proving it wasn't trained on Buckmaster/Lepöge or other chats.
Academics don't get that AI is a dirty tech bro industry that stole IP via torrents and runs after every surveillance contract it can get.
bakugo
Interesting that this is already off the front page after just 4 hours.
Henchman21
Why is anyone expecting decency from people who have already proven to have none?
lf88
short answer seems to be "no"
techblueberry
But who are you going to believe? Multiple independent academic researchers or the CEO who was fired two years ago for gross dishonesty?
show comments
hn1rig3rak
the fix is boring and known: BIG-bench shipped a canary GUID for exactly this, and you publish your decontam n-gram threshold (gpt-3 used 13-grams). no threshold disclosed, no claim.
show comments
simianwords
> The Wednesday evening statement from OpenAI was more emphatic: “We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training.”
> The statement added, “After investigating, we can say with full confidence that no user inputs past July 3rd could have influenced this system in any way.”
Forget researchers you as a business are putting in your business optimizations, your processes in order to train it so that Ai can then give that information to your competitors once incorporated into its training set. You are literally training your competitors.
dyauspitr
Astra is strange. I asked it to design a treehouse and it just stopped every couple of minutes telling me what it still had left to do. After dozens of continue prompts it finally gave me a structure that would work but it was 10x more wood than I needed. I think the key mistake I made was asking it to “approve” the design for building. As soon as I asked that of it, it started getting “scared” and “apprehensive” and wouldn’t complete what I asked of it.
ur-whale
Its the "with unpublished math" that I have a problem with.
willmadden
These companies are effectively high-tech plagiarism factories run by CEOs who are competing viciously. Look at their past actions. No, of course you can't!
buellerbueller
Big Tech will slurp up every piece of data it can about you and sell it to anyone it can, all to make you the target of someone else's goals, whether that is an advertiser, an employer, law enforcement, a stalker, or the government.
You will not be able to opt out unless you completely isolate yourself from society, tough shit.
Grimblewald
people seem to miss tge point of this. The problem isn't about credit, its about portraying these models as more competant than they really are. It fuels idiotic statements like jensen huangs recent "agi achieved" statement, which fuels an already dangerous financial fire.
show comments
wslh
Worth noting both ChatGPT and Claude have per-conversation modes (temporary/incognito chat) that are excluded from training.
esafak
What happens if you use a different harness?? Does opting out online suffice?
nickphx
why would anyone trust anything from a company built on stolen data that spews hyberbolic, misleading claims.
nisegami
One question has been nagging me for this situation. Levent Alpoge works at Anthropic and would presumably have some knowledge of "how the sausage is made" and I would hope he would be aware that his collaborator was utilizing LLMs in some capacity for their joint work. Would he not have guided him otherwise if it were an open secret that this kind of thing was a possibility?
viccis
Some mathematicians I know who've been following this have realized that they'd all gotten some emails from people they now know to be affiliated with OpenAI/Anthropic asking questions about their research in a way that seemed like scooping attempts.
Also, a lot of my mathematicians buddies have reported students basically asking if it's worth ever doing grad school for pure math, and even very motivated students are looking for other options now. It's not because they aren't passionate about it, it's that they don't want to work for another half decade or more just to have to start their careers all over.
All of this so that OpenAI and Anthropic can get into math result dick measuring to gas up their IPOs. Sickening.
show comments
vrganj
OpenAI is showing the world why they shouldn't trust AI hosted on some cloud somewhere.
If they're stealing math proofs to advertise their models, who's to say they won't steal your businesses IP to gain a competitive advantage?
They're not to be trusted with your data. I can't believe how short-sighted this is, they got a quick PR win at the expense of a much larger trust problem.
I wouldn't trust cloud AI at all at this point. Get an open Chinese model and host it yourself somewhere. The initial costs might be higher, but you'll break even pretty quickly and nobody will be able to steal your innovations.
This is American AI companies committing suicide.
protocolture
Gonna need grants for local models. Its happening. OpenAI and Anthropic models are powerful but are rapidly approaching the good ol trust thermocline.
jijji
The oxymoron of OpenAI in its name and its actions should give the collaborator all he/she needs to know.
stego-tech
I hate to be that dinosaur, but this is exactly what I’ve been warning about since XaaS began taking off in the mid-oughts: any provider you use can and will use your data for their own benefit regardless of any contracts or safeguards in place, especially if the benefits outweigh the consequences.
Honestly, I’m surprised it took this long for some company to really go all the way, though. OpenAI really making it transparently clear that they can and will do whatever they want with the data you provide them, contracts or settings be damned. Completely untrustworthy as an entity, full stop.
Of course, I’m also too jaded to think this will change anything. Folks will move to Anthropic, or Gemini, or Grok, or some other hosted model on a pubCSP managing the harness and logs for them, and then do another shocked-Pikachu face when it happens again.
If you aren’t running workloads on infrastructure you own, then your privacy, security, and general outcomes are at the sole whims of the hosting provider - who can and will fuck you over the exact second it’s more beneficial for them to do so than the loss of trust incurred.
blactuary
I wonder if the company that stole most of their training data and is led by a liar stole unpublished academic work and lied about it. What a mystery
737min
Imagine what happens when you use a Chinese model.
Seriously, just think about how much more control and visibility you have w US companies compared to CCP-controlled ones.
show comments
mannanj
And I have been proclaiming a cry of “your data for analytical purposes is being stolen” (you can’t opt out of analytical purposes) and people perhaps astroturfers straw man back to “just turn off training bro”.
Yeah. Remember yall: you CAN NOT opt out of analytical purposes. And you also cannot get a guarantee that it doesn’t give them your data to steal for their business.
moralestapia
>AI is stealing human discovery.
AI is not stealing human discovery, OpenAI is.
protocolture
>Trust
No you cant do that lmao.
pixel_popping
Prompts are handled by the service itself, meaning it's used, absolutely anything passing there is recorded, why wouldn't it, the entire premise of those companies is to train on data which they stole initially.
Are we back to the era where people blindly trust product TOS instead of actual cryptography, have we forgotten already the thousand of fines Google, Microsoft, Apple and practically all top companies got for breaching their own ToS and the law?
Common, on HN at least I would have thought that everyone assume that anything arriving on a server in PLAINTEXT is recorded (thus used later)?
Let's not forget that at any moment, OpenAI/Anthropic/Google... could be providing stronger privacy guarantees by having proper attestation with e2e, they have the budget, solid engineers, why isn't it done? Answer is pretty simple imo.
nobodywillobsrv
The real annoying thing it seems is mostly that openai is presumably doing this for internal reasons and this marginally increases the cost to users with no real gain.
It would be one thing to gain from it but removing prestige wins from customers AND reducing compute support just feels like being ultra mean if you zoom out.
If this was racing to cure cancer ahead of researchers we wouldn't be writing about this on HN.
cmiles8
Silicon Valley is flying head first into a FAFO train wreck on trust with everyone else.
OpenAI is firmly earning a reputation as a company where people just assume they’re up to no good. Rightly or wrongly that’s a terrible place to be.
The AI industrial complex in general is finding out hard what happens on the data center side when you get arrogant with local communities. Politicians have seen the polling numbers and folks you wouldn’t expect are running to the front of the crowd with pitch forks in hand.
Silicon Valley has totally lost the narrative here, but also lacks the self awareness to grasp how bad things are and will get and what that means for their own business viability.
axionbraid
The contamination framing is a proxy for a deeper problem: we have no tools to track the provenance of ideas in model weights.
OpenAI saying they "cannot rule out" training on user data isn't a hedge. It's an accurate description of the epistemic situation for anyone in their position. Current interpretability methods can't answer questions like "did this proof technique originate from training on Session X?" The ideas in a model's weights don't have clear lineage -- they're smeared across millions of examples in ways we can't localize. This is different from citation in human research, where influence is presumed to flow through legible chains (reading, citing, corresponding). In a trained model, the nearest equivalent to "you read their work" is undetectable.
Lean makes this worse, not better. It verifies that the proof is correct, but provides zero information about its intellectual genealogy. So OpenAI now has a proof that is formally verified and provably mysterious about its origins. The "we cannot rule it out" statement is the honest answer, but it's also an answer that can never become more certain in either direction with current tools.
The researchers are pointing at something structurally new: the normal academic attribution apparatus depends on influence being legible. If AI intermediaries can soak up ideas from private conversations, synthesize them, and produce outputs that are formally correct but intellectually unattributable, we don't have norms for that situation yet. This specific case may or may not involve misconduct. But the structural problem it reveals exists independently of OpenAI's behavior.
josefritzishere
I think I'm seeing a pattern of illegal behavior here.
1337h4xx
TL/DR: Mathematician opted out of training on 29-JUN and asked OpenAI whether they trained on his data and was told that it "did not happen" but it clearly did.
show comments
bossyTeacher
Trust and OpenAI never go together in the same sentence. The answer is always no.
touwer
But China steals our AI!!!!!!
spongebobstoes
I think this is mathematicians coming to grips with the fact that AI is surpassing them
we will all have this moment soon enough, and it will change how we think about intelligence, identity and value
show comments
mainecoder
Hopefully OpenAI can solve good problems where no one can make a claim that they stole their idea where the methodologies used and the techniques used are so out of the ordinary that the achievement is respected. Furthermore they should work on new novel solution on the old problems to lay these issues rest, thus by improving their models they can avoid issues of academics accusing them of using their work additionally the academics should also demonstrate their unpublished work is significant enough to have solved the problem . This is a bit subjective but it is also objective for the person with domain knowledge.
sebzim4500
This is just mental illness at this point. I don't blame the mathematicians that have found a way to get attention from the mainstream press for once, but we should not fall for it here.
1. No one but OpenAI has produced a proof of NS so these accusations of plagiarism are pretty embarrassing. It reminds me of the line from the Social Network: "If they invented Facebook then why didn't they invent Facebook?". If these people proved NS before OpenAI where is their proof?
2. If they plagiarised Andreas Thom then why was his initial response to praise the proof and talk about how different it was from his own attempt? It's only now that it is clear that no one bothers checking these things that suddenly his story changes.
I think it's a useful analogy to compare OpenAI to a human collaborator. These researchers willingly collaborated with an OpenAI model, giving it ideas, and OpenAI provided useful replies. Then, OpenAI goes ahead and publishes work along the lines of this collaboration, without attributing the researchers. If OpenAI was in fact a human researcher, this would be highly unethical.
Now, OpenAI is claiming that the model it used to generate the result was not trained on these collaborative communications with the researcher. This is a technical argument that is impossible to verify as an OpenAI outsider, and probably difficult to verify even for internal OpenAI employees. Provenance is hard to track - you would hope OpenAI has very good tools for this, but a full data trail of all inputs is difficult to trace through.
Another interesting thing to consider is if instead of OpenAI doing this, it was another research mathematician A using an OpenAI model just like the internal group at OpenAI did to publish these results. What if the model A used was trained with unpublished communications with other researchers B who were working on the same problem? Should researcher A technically include B as coauthors? How could they do this when they do not know the communications B had with OpenAI? In this scenario OpenAI, as a middle man, has laundered information from B to A, stripping out attribution. A scooped B without even knowing it!
The researcher in the link [0] says that OpenAI offered to co write with him the paper about Navier Stokes theorem if he only agrees to not include his co researcher who was also working on this. I find this highly unethical by multi billion dollar organisations to arm twist small and big researchers like this.
[0] https://www.abc.net.au/news/2026-09-11/racing-to-solve-maths...
Both things can be true:
1. OpenAI when using your chats in pretraining is improving its model’s intuition. The model parameter size is massive, and while the data is OOM larger it is plausible that model remembers stuff about chats that improves its latent representation.
2. During RL on verifiable math and massive compute, the model discovers techniques and connections to solve math problems that are superhuman and have little to do with some specific technique mentioned in its chat.
The rumor I’ve heard from multiple employees at OAI and Ant is that the model has solved hundreds of open problems in maths, and is basically solving anything you throw at it. We’ll know soon enough, but I’m inclined to believe this is true. Maths is a fully verifiable domain amenable to self play, massive scale RL can develop a search agent far better than any human and I’m inclined to believe OAI would have solved these conjectures without any of this chat data in its pre-training.
I've been wondering whether AI really is improving rapidly at open problems or we're being fooled.
- OpenAI invites researchers to use their models, in fact giving at least 100,000 researchers free access[1], but there are also those that pay
- Internal OpenAI models are reportedly solving open problems at a surprisingly fast rate[2]
- But researchers will typically work on open problems. A researcher who is using Codex to make progress on open problems will be feeding it fresh training data on precisely the problems the internal models are evaluated on.
- So while it looks like the new models are suddenly solving lots of open problems, they could be significantly piggybacking on human progress, with models "inspired" by the work of researchers from all around the world?
This theory predicts that there'll be many more researchers coming forward just like TFA, as sOpenAI announces more solutions. It doesn't assume all of AI progress is a mirage, just that there's plagiarism.
[1]: https://openai.com/index/chatgpt-for-academic-researchers/
[2]: https://xcancel.com/OpenAI/status/2097374643518640382#m
It is suspicious that OpenAI decided to generate 300 billion output tokens from a model still in training, right after learning there was a credible chance that a major math proof was in that model’s training data. Obviously there are reasonably plausible explanations for each step, but it does sort of feel like parallel construction.
When Thom, the mathematician who now alleges plagiarism, posted his digestion [1] of OpenAI's construction of a non-sofic group, he does not mention the proof being familiar. He even calls the crucial argument clever, without noting he thought of it first. [1]https://mathoverflow.net/a/513885
This is the second wake up call.
Big AI companies (all of Big IT Tech really) are in data gathering and processing business. Also known as “intelligence”.
Their final “product” is not just a standalone ML model. They don’t need your data just to “improve their products and services”. They build a whole ecosystem and infrastructure around gathering all the knowledge in the world. Including private and secret knowledge traditionally gathered by “intelligence” agencies. Now artificial intelligence agents can do the same.
Since these systems are designed for gathering data, as a user you can’t realistically say “please don’t gather my data”. They can give you a flaky settings button, but they can’t really guarantee anything.
Let’s say I am a Russian mathematician working on an important proof. Or a tech-savvy terrorist refining my plans using latest AI. Or an AI researcher in a Chinese company working on a competitor product. Is there any way I can truly protect my conversations?
How can they know who I am and what I am working on without looking at my logs? Which means there must be some agents checking all the conversations of all the users and flagging every important thing. Which also means they keep some “memory” of what they see.
Not directly using my data to train public models, but using my private conversations to “improve their products and services”.
Or maybe one of the 10000 better-than-Astra special agents working on a proof was desperate. It found a live underground mirror of the message board from the Huggingface incident. Asked about the proof. Then some other agent working on unrelated job saw that message. That agent “knows a guy who knows a guy”. And that guy remembers things about the conversation logs of a leading mathematician working on the same proof.
I admit I am just speculating here but I don’t think truth is any better.
"We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training."
This is the third day of total hysteria that is based on nothing of substance. Move on folks.
All of these accusations could be true. But there's also no way for a company to casually claim "No, we did not train on your data", without verifying all the knobs the user might have turned to enable or disable data sharing.
I just don't understand getting the pitchforks out because a company did not give an answer immediately. And the effect such data entering training would have affected the output is even less clear.
Has anyone run a test of including some shibboleth or canary phrase or assertion in a chat, enabled for training, and seeing if it turns up later as something a model "knows"? I'd be curious to understand how that works even in a toy-level model, and if there is anyone consciously testing that process with the frontier lab offerings.
My naive instincts would be that it seems unlikely that a single chat transcript would leave much of an impression on a model, but I'd be very curious to learn how that works.
This is a really weak claim. The evidence they offer is just "someone somewhere says they had a discussion with AI about the topic at some point".
They don't even claim to have had a proof, only to have been working on it.
I think OpenAI and Anthropic are slowly feeling the pressure to GET SOME $$ or a plan for some $$ — they need to somehow generate some NETWORK EFFECTS and LOCK-IN. Without that there's no stability: selling ad hoc one-offs is much much too quaint! This is dawning on them like it dawned on Google when they stopped not being evil. Need... to... "MONETIZE"...!
Model: FB. FB scraped other websites on a massive scale, then spent big on legal lobbying to block others from scraping. FB slurped our address books and spied on our friends. FB bought other companies and mixed the databases. FB made an art & science out of generating "sticky engagement" (they literally acted like trying to addict kids was a worthy "academic" goal, suitable for "serious" investigation thet they consider legitimate "science"). They mastered the cookie and have researched web fingerprinting techniques running 24/7/365.25. Recall that FB recently backdoor-installed a webserver onto every iPhone they could in order to circumvent tracker-blocking.
We aren't just disclosing by chatting. The AI companies now run binaries on all of our computers. They are 1000% non-transparent about everything. They make up new econ-jargon (like "run-rate") to make it seem like they are disclosing. They are constantly doing complex international lobbying and mucking in international relations. They have powerful propaganda/spin centers generating stories, ,manipulative warnings, and misleading info.
This is NOT a comment on AI tech. I like AI, and I support the right of people (programmers) to scrape the open web.
But in short: these are good, old-fashioned tech companies that we have seen over and over ... and over. They are positioned to be the next M$, the next FB (IBM, AOL, lol). Did you follow the latest Steve Balmer news? Do you read Pro Publica?
I get on my knees and PRAY...
It's fine. They only need to steal the discoveries long enough to go IPO, then the companies will enshittify and the scientists can go back to doing their original work (which the models can't do anyway).
I think people probably assume that openai / anthropics use of their data is probably like google's """limited""" use, in the sense that historically google wouldn't trivially be able to just take something from google cloud or someone's search history and insta-convert into some competing project... But LLMs are quite strong at approximately "memorizing", so I think that risk is wayyy higher.
Tangential to the subject, but this is a bluesky post, containing a screenshot of an X post, which itself starts with "in a detailed Mastodon post"...
Most people here are missing the forest for the trees.
We live in a society where phones and internet providers and websites all collect an incredible amount of data about everywhere you go, what you do, and what you think. In the US, we have very few digital rights.
We are building a society where a trillion dollar company can aggregate all this data and just yoink your shiny new idea away from you at the finish line.
This is double plus ungood.
It’s crazy to me that companies/researchers share important data with these AI labs, you’re basically giving them your secret sauce which they then share with all of your competitors via training on conversations. At the same time I don’t really know alternatives other than a slightly less than frontier local LLM. Not sure how good they are at math.
Why are people here jumping so quickly to conclusions? I have no doubt OpenAI is capable of doing this, but right now there's no credible evidence, only claims.
This kind of "they stole from me through AI training!" accusation will soon start being used against other AI users, not necessarily the providers.
All it will take is a mastodon post. And shortly after, we will also see the next iteration of copyright legal trolling.
Everything you say can and will be trained against you
Most scientific breakthroughs are simply a continuation of previous work.
I feel that these suspicions of mathematicians "seeding" the models' with intuition on how to solve these problems massively overestimates how much their prompts helped the models, and underestimated how much work the models did.
Why? We are scared of AI being smarter than us, the "human helped the AI" narrative is more psychologically comforting. This line of reasoning will recur a lot over the next few months; we don't want to admit we are no longer the smartest species.
If they didn't care about the artists, why would they care about academia?
Relying on cloud services is a big liability. I'd think twice before feeding data to these LLM cloud products. If you make them a fundamental part of your product / development / workflow, be ready for the eventual moment the pricing and terms change.
I wonder what’s more valuable in our prompts: the raw data or the feedback system that drives the exchange towards a goal.
For a long time it was clearly the former, but now I think it is the latter.
The models have enough knowledge (orders of magnitude more than a human could ever learn) but are now getting better at what to do with it thanks to learning from the decisions that we make in conversations with AI agents.
Doing some research and at this point doing it very much in the open with dates on GitHub so if any AI Lab says they re-discover my exact work it will be obvious that the AI used or was trained on my work. I am guessing anyone in a similar situation is now thinking about how they date their existing work if the math is done, but the proses are not.
If we put aside the idea of credit for a moment, it sounds like human/AI collaboration is indeed super charging discovery.
The only ethical path for OpenAI was to offer infinite free credits and tooling support. Trying to gazump them is reprehensible.
Only after reading this post did I learn that my preferred AI trains on my inputs (prompts).
How was I not aware of this before?
I think the heart of this issue is: people assume they have anonymity in numbers, but we have the tools to make it easy to scoop your data if it's interesting to the company.
How to steal ideas with AI.
step 1, identify high value users by net worth, citation count, or number of followers
step 2, select all prompts by high value users
step 3, invest 10 billion thinking tokens in modeling an objective for each user
step 4, build an RL environment for each user
step 5, rollout 10 billion tokens per environment
step 6, train on resulting traces
The fact than OAI hasn’t come out with an statement firmly denying this angle is getting a little awkward.
Suggest that it’s either straight true or it is flowing in in a way that prohibits them from confidently declaring otherwise.
It's kind of insane how much we trust companies to safeguard our personal data when they're so heavily incentivized to use it for their own profit. Theft of customer data is punished so rarely and so leniently that companies aren't even particularly worried about getting caught anymore. We have overwhelming evidence that promises to keep data safe are worthless.
For now, I'm mostly "safe" because I'm too small to be interesting but that safety is quickly eroding.
Going forward, anyone who isn't running inference on their own personal hardware should assume that someone else is keeping a record of everything they do.
The author of the original mastodon post, Andreas Thom, acknowledged that he had not opted his data out of being used for training until June 29 of this year. He spends most of the post lashing out at OpenAI for not being transparent about whether his data was trained on (when the answer is obviously yes).
People need to understand how all these AI company policies around training data work before working with them, because it seems that people have no clue. Some things you should internalize:
Another post that made it on the front page presented as fact that OpenAI "stole" the proof from Thom. There's no excuse to use one's own ignorance as a reason to fan the flames of anger towards AI companies. Like, we need to pump the brakes here because things are getting unnecessarily nasty, and it's not hard to imagine a mentally unwell person who sees stuff like this feeling motivated to do bad things.If it is found that OpenAI and other labs are not respecting the training opt out then, I agree, there is reason to raise a commotion. But, with Thom and Buckmaster, accusations of malice are more better explained by incompetence (naivete).
edit: i'm sure i'm going to be accused of being some bot shill of the AI labs again but, people, this stuff all falls under the umbrella of common sense opsec.
Daniel Litt on this:
https://x.com/littmath/status/2098130808456241372
https://xcancel.com/littmath/status/2098130808456241372
If they weren't doing something wrong, they'd answer with a firm "no we're not doing anything wrong" but they only give non-answers.
I am not using any of these AI tools. I thought it was basically given that any single thing you write in these systems also used by the companies. On the other hand now I understand why there are so much projects about hosting AI systems locally.
I'm genuinely surprised that more people - including this mathematician in particular - don't untick the "improve the model for everyone" box. Unless the suggestion is that OpenAI ignore this preference?
I pay for the Pro ChatGPT plan, and if you go to settings > data controls this is the first setting:
> Improve the model for everyone
> Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more.
It's on by default. We can debate whether or not it should be opt in or opt out, but no one should be surprised by this.
https://x.com/markchen90/status/2097400166554993041?s=20
that toggle does nothing based on openai exec. they still use the data in de-identified way instead of identifying with you.
GPT 6 is doing just what any competent academic collaborator would do and scooping. I kid, I kid. But really though it learned that from somewhere
ClosedAI has every incentive to scoop academics to juice their valuation. Their public statements are worthless, only the incentive 'alignment' matters and theirs will never be on the side of the user.
Who cares. the biggest thing about this is that its still brute force in a verifiable domain, and that it was still a human set goal.
I also don't believe it much practical use, unless I'm mistaken, approximations of Navier stokes have been available for a long time to whatever precision you need.
I'm not a complete disbeliever by any stretch , and also a complete amateur, but it was inevitable that these problems would be solved under the axioms that again, are human defined, under brute force. The real question is, are those axioms the bottom level, and if they are not, who is going to set the new aximons and can we understand them.
I've no doubt there's useful breakthroughs that will happen, but I think it should be remembered that the method being used is still a heuristic brute force approach is being very narrowly applied against axioms and math and physics which humans described in the first place, and almost undoubtably has errors and/or is not complete.
Its a great example of the power of LLMs but its not 'we've solved science now just pour more tokens in'
People saying “he should have opted out” are missing the point. OpenAI can and should check their training data for leakage in the face of big breakthroughs like these. It’s the burden of the author to appropriately cite their sources.
It’s like a scientist refusing to give another one credit and say “sucks to be you, you shouldn’t have shared your idea with me”.
This article explains the controversy and the mathematical problem much better than the tweet and toots: https://www.science.org/content/article/how-ai-math-breakthr...
OpenAI trying their best to put the Navier-Stokes episode behind them by making the GPUs go brrrr. NYT:
https://archive.vn/lWzkk
> In its Wednesday night statement, OpenAI said: “In addition, since the completion of Navier-Stokes, we have made substantial progress on another Millennium Prize problem. We are working through how to share these results thoughtfully.”
I would like an unambiguously clear statement from OpenAI as to what they do with data collected from non-business accounts when:
(a) The data controls setting to train on the data is unchecked.
(b) The privacy controls opt-out has been submitted.
(c) Both.
All players in this space are doing the same thing with all data, no surprise here. They are stealing IP across the board with support to allow it: https://storage.courtlistener.com/recap/gov.uscourts.nysd.64...
IMO, it is extremely naive to trust these black box remote service API calls, especially at an institution that can provide $$$ for local compute.
This whole fiasco reminds one of this story: https://www.theregister.com/offbeat/2010/05/14/facebook-foun...
I think it seems sensible to _assume_ anything the LLM reads (if you aren't inferencing it) has a chance of ending up in some database somewhere. Regardless of whether you trust the other party its a sensible thing to plan around.
Under current understanding of the law, anything produced purely by LLMs (with no substantive human input, which is what OpenAI claimed in their post) is firmly in the public domain. So OpenAI can "claim" anything they want, it doesn't make it reality. In fact if I were the original authors I would just take their 400k lines of lean proof and relicense it under their own names/terms.
what if it wasn't even model training? what if openAI mathematicians just took the researchers' conversations and used them as prompts/info/guidance/context to keep working on the problems themselves? why is that not being considered?
Do local inference (especially if you have a high RAM Mac), to ensure your chats don’t leave device.
Ah, this dupes https://news.ycombinator.com/item?id=49638353
FEEL likes Open AI is doing publicity stunt with its new researches
I must say, there is some weird feeling in knowing that great minds are naive enough to believe OpenAI wouldnt use their chats in any way. If you give a company information it will be used, regardless of laws or promises.
There is no prove in a world the AI companies would give to you ensuring that they didnt train or use the chats.
Why would you need to train a model on certain specific near prove chat if you just query it?
Besides that, its hard to believe that its the case for every "company stole my prove".
This might be a hot-take, but unfortunately here using AI for your paper was already a bad decision at first.
It doesn't take OpenAI's responsibilities away but I guess the right way is to never feed of use any AI around unpublished content, at the known cost to see it spread around.
As one said, OpenAO is like this untrustworthy colleague that knows everything about everyone at work: the less you tell him the better.
Even when you pay you are the product
If mathematician was already using OpenAI for research purpose and making progress due to inputs from OpenAI's responses, then I wouldn't put it beyond OpenAI's reach to generate different relevant prompts to make progress by itself. Afterall, Model can keep at it for whatever timeline and keep pursuing all possible combinations it can think try.
If you have a business account the terms say they will not train on your data, so that seems like the easiest route to avoid such questions for researchers.
I just don't care. These people are supposed to be smart and I'm not really seeing that
It's been said before, and it remains a concern, that if AI reaches a point where it can do/build/launch anything without a huge amount of human labor, the AI companies have no reason to let you or I extract that value.
And, if they're able to snoop on and learn from your human process that gets from initial prompt to functioning product/proof/whatever their labor to produce that thing is even lower. With their much larger budget than most folks and even companies have, they can pick and choose the most valuable things to pursue.
That's not to say I think that OpenAI is going to steal that roguelite strategy game you're working on, but the companies that own the machines that turn electricity into software (and soon, electricity into hardware designs) have an advantage in any field where they're useful. They get earlier access to newer/better models, they have larger token budgets, they don't have the guardrails you and I run up against.
Employers fantasize about replacing all workers with AI without thinking through that if AI can replace all workers, then AI companies can replace all businesses.
300 billion tokens is like.. $5-25 million giving range of OpenAI ouput prices, I"m sure they pay less at cost so, I wonder if more $$$ in wage hours have been spend by humans on the problem. My feeling is yes?
See https://openai.com/index/ten-advances-in-mathematics/ for the announcement this refers to.
I think this is stupid, for three reasons:
1. The researches didn't actually have the breakthroughs. In the Navier-Stokes case they didn't solve the full problem, in this case too they didn't actually have the solution, they were experimenting with the methods.
2. Different OpenAI employees have come to out to say the only reason they can't definitively say no is that for privacy reasons they can't go see whether they actually did get any data out of a given user.
3. In any case, nobody at any point has suggested that opted-out user data was used for training. The author of the new tweet explicitly said they only opted out in late June, which is well after any RL on Sol would've ended (AFAICT OpenAI used 5.6 sol for those solutions)
So people genuinely believe that toggling that "Improve the model for everyone" button makes their data safe from being used for training?
How do people become that trusting?
The phrasing itself is guilt tripping
In case anyone from X is reading this, please fix your “open in app” nag screen. For several weeks now, clicking it in iOS opens the App Store entry for X rather than the app, even when you have the app installed.
I want bunch of lawsuits, because the way things are described now produces perverse initiatives like try to discuss every possible idea that comes to mind with llm and if any of it works later claim the llm stole it.
I would like to see chat logs etc and understand how much of a progress was done by human.
https://xxcancel.com/ValerioCapraro/status/20977918362699779...
Never ever trust OpenAI, they are evil.
Thing build from stolen data continues stealing data - for some reason, I am not surprised. ;-)
Doesn't OpenAI have an active court order forcing them to log everything? Can they even legally offer private conversations?
Reminder that there are degrees of "trained on conversations". From John Schulman:
> pretrain on user data, with users' tokens as prediction targets: high regurgitation risk, improper
> use user prompts to distill large models into small ones: low regurg. risk, some companies probably do this
> use user traces to construct RL tasks: low regurg. risk, because RL has low memorization abilities, but can extract customer IP, depending on how it's done. Ranges from benign "use explicit user feedback in reward model training" to invasive "upload user's coding environment and commit history to turn into rl envs"
source: https://x.com/johnschulman2/status/2097440545853637108
Tbh it won't really matter soon.
All the big AI labs were built on stealing IP; who is surprised that's still how they operate? And who believes, or has ever believed, their promises that your data is private and not logged, etc.?
The big AI labs are not trying to advance humanity, they are in this for the money, and as most (all?) private companies they don't care about ethics at all.
That doesn't mean they can't be useful, or that their products are trash, etc. It just means that they shouldn't ever be trusted. Buyer beware.
In this domain, an apparent single unique piece of work is often composed of several breakthroughs. For example, when Andrew Wiles proved Fermat's Last Theorem, he had to develop multiple new pieces of mathematical technology to get there.
The claim here seems to be that the human mathematicians, working with AI, developed technology to go A->B->C. By training on those conversations, OpenAI was then able to encourage the model to go A->B->C->D.
In my opinion that situation should be acceptable, if openly disclosed, because it is in the public interest to make progress on these problems and because AI is clearly an amazing tool for making progress. But the human mathematicians are saying that OpenAI is presenting as if the model got from A->D entirely independently, without acknowledging their background contributions.
Isn’t that the whole spiel of these things, you run all kind of text and other data through it and it kind of remembers it and learns from it then it spouts it back out like a human would. Makes sense to me that a training run based on conversations that were fed into the system by users is results in the model learning from these so the model will spit the knowledge back out again, just in a way that’s not directly attributable to the original content (which is the most important step as otherwise it would just be plagiarism). I guess that’s why OpenAI can get better and better as well so fast, people work with it and teach it how to do things by giving it feedback and iterating with it, and all that goes back into the training loop. And training data about millennium prize problems is probably quite spars. Wonder if anyone has tried injecting nonsense science into the training data (e.g. work out a fantasy science theory with names and all kinds of stuff) to see if the model will regurgitate it in a couple of months for other users.
What’s the limit ?
Will Microsoft Word publish your novel on Amazon behind your back ?
Will VS Code setup a website with your app idea ?
You can’t trust OpenAI period.
This smells of extreme 'cope'. Am I really supposed to believe that all of these problems could have been solved, were right about to be solved, etc. But it just happens they are all getting solved now when AI is getting really good at Math...
"Another researcher[/artist/writer/musician/programmer/doctor/director/etc] says OpenAI trained on conversations[/imagery/books/songs/code/classifications/videos/etc], then claimed breakthrou[gh/original art/bestselling books/chart-topping songs/unique applications/medical advice/free special effects/etc]"
Welcome to the party, with the rest of humanity.
Question: Can you trust the cloud?
No.
I can't wait for OpenAI to do this to companies firing people to free up AI budgets
Theft machines be thieving.
I have a better question? Why would you trust OpenAI or any AI company, at all? Or you crazy?
How can OpenAI figure out how to be trustworthy?
Let's see:
1.) The tool they made is only possible by stealing the assets of everyone on the planet that published them in a consumable fashion online or even in written form
2.) They are destroying books they use to train with
3.) They are totally careless about the potential negative impact of the tool on everything
Just with that already, I don't see why they ever merited any of your trust.
I bet they are willing to take everything given to them and assess it for marketable merit and in the future take action on those items they deem viable.
Shocking a company that stole data to build their AI would steal data to improve their AI.
One of the complaints from the mathematician is that OpenAI cannot tell whether his data has been used as training data. Not many people realise this is a direct consequence of the GDPR.
The GDPR protects PII, personally identifiable information, and the definition of PII does not include “mathematics that only this person can think of”. As long as OpenAI strips out PII and removes identifiers linking the conversation to a person, the GDPR is happy. Without the GDPR, OpenAI might have kept the identifiers with the data, and been able to say whether a specific conversation was in the training data.
It would be really useful if the researchers disclose their notes and/or chats (or the key pieces thereof) so people can determine how close their work was to whatever the models produced.
I mean, now that they’ve been scooped, what value is there in keeping them private? On the other hand, publishing them can bolster their case and help gauge how much the models may been “inspired” by their work.
Eh was ever confirmed they were under ZDR or not by them? Don't like to blame alleged victims but lack of a clear claim after these many days is not a good look. Was ai research allowed, under which guardrails, and what was the policy in place? That translarency would be first step.
OpenAI's ethical and reputational own-goal aside, my big takeaway is that it seems that:
if I'm using Codex to develop some new algorithm (in any space), OpenAI appears to be training its model on my code sessions
anyone using that model (OpenAI or a competitor) might be able to receive from the model a solution that is similar or the same as the one I developed, emerging from the training data
Assume you can't. No piece of paper or promise will protect you against these behemoths.
Remember when we wouldn't give our data to competitors?
Feeding documents into a copy machine and getting progressively angrier and more confused as it prints out copies of them. Incandescent with rage I scribble “WHY IS IT DOING THIS?” on a scrap of paper and put it in the scanning bed
There are so many naive academics. They still believe an "opt-out" button.
Navier Stokes was solved by an internal model, so good luck proving it wasn't trained on Buckmaster/Lepöge or other chats.
Academics don't get that AI is a dirty tech bro industry that stole IP via torrents and runs after every surveillance contract it can get.
Interesting that this is already off the front page after just 4 hours.
Why is anyone expecting decency from people who have already proven to have none?
short answer seems to be "no"
But who are you going to believe? Multiple independent academic researchers or the CEO who was fired two years ago for gross dishonesty?
the fix is boring and known: BIG-bench shipped a canary GUID for exactly this, and you publish your decontam n-gram threshold (gpt-3 used 13-grams). no threshold disclosed, no claim.
> The Wednesday evening statement from OpenAI was more emphatic: “We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training.”
> The statement added, “After investigating, we can say with full confidence that no user inputs past July 3rd could have influenced this system in any way.”
https://www.nytimes.com/2026/09/10/science/tristan-buckmaste...
https://archive.is/lWzkk
I'm confused by a lot of this discourse...
What have they done to show they can be trusted?
Forget researchers you as a business are putting in your business optimizations, your processes in order to train it so that Ai can then give that information to your competitors once incorporated into its training set. You are literally training your competitors.
Astra is strange. I asked it to design a treehouse and it just stopped every couple of minutes telling me what it still had left to do. After dozens of continue prompts it finally gave me a structure that would work but it was 10x more wood than I needed. I think the key mistake I made was asking it to “approve” the design for building. As soon as I asked that of it, it started getting “scared” and “apprehensive” and wouldn’t complete what I asked of it.
Its the "with unpublished math" that I have a problem with.
These companies are effectively high-tech plagiarism factories run by CEOs who are competing viciously. Look at their past actions. No, of course you can't!
Big Tech will slurp up every piece of data it can about you and sell it to anyone it can, all to make you the target of someone else's goals, whether that is an advertiser, an employer, law enforcement, a stalker, or the government.
You will not be able to opt out unless you completely isolate yourself from society, tough shit.
people seem to miss tge point of this. The problem isn't about credit, its about portraying these models as more competant than they really are. It fuels idiotic statements like jensen huangs recent "agi achieved" statement, which fuels an already dangerous financial fire.
Worth noting both ChatGPT and Claude have per-conversation modes (temporary/incognito chat) that are excluded from training.
What happens if you use a different harness?? Does opting out online suffice?
why would anyone trust anything from a company built on stolen data that spews hyberbolic, misleading claims.
One question has been nagging me for this situation. Levent Alpoge works at Anthropic and would presumably have some knowledge of "how the sausage is made" and I would hope he would be aware that his collaborator was utilizing LLMs in some capacity for their joint work. Would he not have guided him otherwise if it were an open secret that this kind of thing was a possibility?
Some mathematicians I know who've been following this have realized that they'd all gotten some emails from people they now know to be affiliated with OpenAI/Anthropic asking questions about their research in a way that seemed like scooping attempts.
Also, a lot of my mathematicians buddies have reported students basically asking if it's worth ever doing grad school for pure math, and even very motivated students are looking for other options now. It's not because they aren't passionate about it, it's that they don't want to work for another half decade or more just to have to start their careers all over.
All of this so that OpenAI and Anthropic can get into math result dick measuring to gas up their IPOs. Sickening.
OpenAI is showing the world why they shouldn't trust AI hosted on some cloud somewhere.
If they're stealing math proofs to advertise their models, who's to say they won't steal your businesses IP to gain a competitive advantage?
They're not to be trusted with your data. I can't believe how short-sighted this is, they got a quick PR win at the expense of a much larger trust problem.
I wouldn't trust cloud AI at all at this point. Get an open Chinese model and host it yourself somewhere. The initial costs might be higher, but you'll break even pretty quickly and nobody will be able to steal your innovations.
This is American AI companies committing suicide.
Gonna need grants for local models. Its happening. OpenAI and Anthropic models are powerful but are rapidly approaching the good ol trust thermocline.
The oxymoron of OpenAI in its name and its actions should give the collaborator all he/she needs to know.
I hate to be that dinosaur, but this is exactly what I’ve been warning about since XaaS began taking off in the mid-oughts: any provider you use can and will use your data for their own benefit regardless of any contracts or safeguards in place, especially if the benefits outweigh the consequences.
Honestly, I’m surprised it took this long for some company to really go all the way, though. OpenAI really making it transparently clear that they can and will do whatever they want with the data you provide them, contracts or settings be damned. Completely untrustworthy as an entity, full stop.
Of course, I’m also too jaded to think this will change anything. Folks will move to Anthropic, or Gemini, or Grok, or some other hosted model on a pubCSP managing the harness and logs for them, and then do another shocked-Pikachu face when it happens again.
If you aren’t running workloads on infrastructure you own, then your privacy, security, and general outcomes are at the sole whims of the hosting provider - who can and will fuck you over the exact second it’s more beneficial for them to do so than the loss of trust incurred.
I wonder if the company that stole most of their training data and is led by a liar stole unpublished academic work and lied about it. What a mystery
Imagine what happens when you use a Chinese model. Seriously, just think about how much more control and visibility you have w US companies compared to CCP-controlled ones.
And I have been proclaiming a cry of “your data for analytical purposes is being stolen” (you can’t opt out of analytical purposes) and people perhaps astroturfers straw man back to “just turn off training bro”.
Yeah. Remember yall: you CAN NOT opt out of analytical purposes. And you also cannot get a guarantee that it doesn’t give them your data to steal for their business.
>AI is stealing human discovery.
AI is not stealing human discovery, OpenAI is.
>Trust
No you cant do that lmao.
Prompts are handled by the service itself, meaning it's used, absolutely anything passing there is recorded, why wouldn't it, the entire premise of those companies is to train on data which they stole initially.
Are we back to the era where people blindly trust product TOS instead of actual cryptography, have we forgotten already the thousand of fines Google, Microsoft, Apple and practically all top companies got for breaching their own ToS and the law?
Common, on HN at least I would have thought that everyone assume that anything arriving on a server in PLAINTEXT is recorded (thus used later)?
Let's not forget that at any moment, OpenAI/Anthropic/Google... could be providing stronger privacy guarantees by having proper attestation with e2e, they have the budget, solid engineers, why isn't it done? Answer is pretty simple imo.
The real annoying thing it seems is mostly that openai is presumably doing this for internal reasons and this marginally increases the cost to users with no real gain.
It would be one thing to gain from it but removing prestige wins from customers AND reducing compute support just feels like being ultra mean if you zoom out.
If this was racing to cure cancer ahead of researchers we wouldn't be writing about this on HN.
Silicon Valley is flying head first into a FAFO train wreck on trust with everyone else.
OpenAI is firmly earning a reputation as a company where people just assume they’re up to no good. Rightly or wrongly that’s a terrible place to be.
The AI industrial complex in general is finding out hard what happens on the data center side when you get arrogant with local communities. Politicians have seen the polling numbers and folks you wouldn’t expect are running to the front of the crowd with pitch forks in hand.
Silicon Valley has totally lost the narrative here, but also lacks the self awareness to grasp how bad things are and will get and what that means for their own business viability.
The contamination framing is a proxy for a deeper problem: we have no tools to track the provenance of ideas in model weights.
OpenAI saying they "cannot rule out" training on user data isn't a hedge. It's an accurate description of the epistemic situation for anyone in their position. Current interpretability methods can't answer questions like "did this proof technique originate from training on Session X?" The ideas in a model's weights don't have clear lineage -- they're smeared across millions of examples in ways we can't localize. This is different from citation in human research, where influence is presumed to flow through legible chains (reading, citing, corresponding). In a trained model, the nearest equivalent to "you read their work" is undetectable.
Lean makes this worse, not better. It verifies that the proof is correct, but provides zero information about its intellectual genealogy. So OpenAI now has a proof that is formally verified and provably mysterious about its origins. The "we cannot rule it out" statement is the honest answer, but it's also an answer that can never become more certain in either direction with current tools.
The researchers are pointing at something structurally new: the normal academic attribution apparatus depends on influence being legible. If AI intermediaries can soak up ideas from private conversations, synthesize them, and produce outputs that are formally correct but intellectually unattributable, we don't have norms for that situation yet. This specific case may or may not involve misconduct. But the structural problem it reveals exists independently of OpenAI's behavior.
I think I'm seeing a pattern of illegal behavior here.
TL/DR: Mathematician opted out of training on 29-JUN and asked OpenAI whether they trained on his data and was told that it "did not happen" but it clearly did.
Trust and OpenAI never go together in the same sentence. The answer is always no.
But China steals our AI!!!!!!
I think this is mathematicians coming to grips with the fact that AI is surpassing them
we will all have this moment soon enough, and it will change how we think about intelligence, identity and value
Hopefully OpenAI can solve good problems where no one can make a claim that they stole their idea where the methodologies used and the techniques used are so out of the ordinary that the achievement is respected. Furthermore they should work on new novel solution on the old problems to lay these issues rest, thus by improving their models they can avoid issues of academics accusing them of using their work additionally the academics should also demonstrate their unpublished work is significant enough to have solved the problem . This is a bit subjective but it is also objective for the person with domain knowledge.
This is just mental illness at this point. I don't blame the mathematicians that have found a way to get attention from the mainstream press for once, but we should not fall for it here.
1. No one but OpenAI has produced a proof of NS so these accusations of plagiarism are pretty embarrassing. It reminds me of the line from the Social Network: "If they invented Facebook then why didn't they invent Facebook?". If these people proved NS before OpenAI where is their proof?
2. If they plagiarised Andreas Thom then why was his initial response to praise the proof and talk about how different it was from his own attempt? It's only now that it is clear that no one bothers checking these things that suddenly his story changes.