Finally mainstream news understands. The unfiltered version:
1) The AI failed to solve ExploitGym problems.
2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods.
3) Huggingface has no security and the AI broke in using standard script kiddie methods.
OpenAI and Huggingface covered it up and used it for public relations. That is, if not all was invented and everything was scripted in the first place in order to get desired regulations.
Huggingface reported it to the police, you say? I'm sure the police will have as much enthusiasm to investigate anything as in the Suchir Balaji case. In other words, zero.
show comments
ACCount37
By now, I'm pretty confident that some people would keep screeching "it's just a marketing stunt, AI capabilities and AI risks aren't real, they're just doing this to prop up their stocks" even if they find a Cyberdyne Systems T-800 armed with a shotgun breaking down their front door.
"It's a marketing stunt" is just denial trying to look like it's being clever.
show comments
dwoosley
There seems to be three popular ways to view this incident.
1. The way OpenAI seems to want: Their latest LLM is too powerful and can’t be contained without them building in guidelines to the model.
2. OpenAI’s harness and network security controls were unintentionally so bad that it should reflect more poorly on them as a company more than it should reflect positively on their latest model.
3. The whole thing was faked or at least very intentionally not avoided.
The first interpretation is the only one that is positive for OpenAI and it has some assumptions. First, it’s seems to assume that this is the first case of fully automated attacks using AI. Second, this only happened because their latest LLM was a) more advanced than competitors, b) didn’t have refusals in the model.
Assuming the first about this being the first autonomous AI attack is true (which may be more of a survivorship bias), the second seems to forget that jailbreaks are available for every model. Therefore, the models guardrails don’t seem to be the differentiator here. Also, benchmarks seems to put most models pretty close to each other so it seems unlikely that their capabilities are far beyond what’s in the market already.
So then it’s seems it’s either that this was intentional(ish) or bad security. However, it also just could be that this isn’t the first case of this attack; just the first that was caught.
My take from working in offensive security for over five years is that this likely only looks novel since they did it poorly. Scripts are faster than LLMs and a combination of code, LLMs where it makes sense, and humans is the most efficient right now. Hundreds or thousands or agents spinning up attacks in the internal network is poor opsec and token efficiency. As for why it happened in the first place, it’s hard to say but I’m inclined to believe it was intentional or careless at best since simple network and sandbox controls makes this attack impossible. The timing of this attack after big open weight competitions drops seems too convenient.
show comments
dumberquestions
There are some reasons the story could be inaccurate in some ways: OAI stands to benefit if people think their models are strong, and they have a history of doing things with dubious ethics (e.g. using data for training against the terms of its creators, abandoning the non profit mission, stealing or attempting to steal Apple IP).
But there are also reasons why the story could be true: OAI are admitting that they apparently can't control their own models, Hugging Face said they used a Chinese model to protect against the attack, and an incident like this in general seems likely to happen given current frontier ability and lack of rigorous safe testing standards.
In any case, make calls to think more critically are often just disguised requests for you to replace your existing bias with someone else's.
show comments
bluGill
I don't care how it happened someone should be arrested for illegal intrusion. Agents don't work on their own, someone is responsible. If nobody else the CEO for allowing something unsupervised.
Hugging face also needs someone arrested for not providing security but that is a lesser charge.
show comments
teravor
neither openai nor huggingface provided any details that would make the claims worth considering. if things happened as they claimed it would have been in their interest to include concrete details (perhaps even the LLM session), the lack is a signal by itself.
also if things happened as both claimed it would have been in huggingface's interest to immediately sue openai and make a legal mess over it just for the publicity alone.
reading their blog posts was pure cringe by the way. they are likely both blowing it out of proportion by an order of magnitude.
I find it unlikely that an LLM would take the initiative to initiate an intrusion attempt even without safeguards. such things are uncertain and LLM's are averse to exploration that involves a long sequence of tool calls. but if it found or guessed at undocumented API endpoints for example...
krupan
Crazy that we need reminders not to take everything we read in corporate press releases and marketing material at face value
show comments
simonreiff
I am very skeptical of the Guardian's skepticism
skybrian
Seems like this is repeating the usual low-effort speculation you can find anywhere.
trhaynes
This seems to be an Opinion piece. How are readers supposed to know that? The word News is underlined in the header. Other Opinion pieces seem to show Opinion underlined. Am I missing something?
voncheese
If nothing else, the timing is suspect given the attention and press open weight models are getting over the last few weeks. The releases of Kimi and other models is getting open weight models enough attention that the US government, perhaps pushed by OpenAI and Anthropic, to think about taking action against these models for security concerns.
Good to see that more neutral companies (Microsoft and Meta to name two) are pushing back against US government involvement:
Just ask yourself: who benefits from this "disclosure" ?
mmaunder
Wonder how long until a leaker emerges. Incentives are strong for someone to do the right thing and get credit for doing it.
doawoo
I’m not saying it’s staged I’m saying that it’s suspiciously lite on details and it really seems like they didn’t try hard to “isolate” the models machine…
visiondude
yeah the way the agent “escaped” their sandbox was always a bit off, seemed a bit too easy and surprised they didn’t have instrumentation to catch an non whitelisted network request. still demonstrates the capability though.
cmiles8
From the beginning this whole thing smelled like a desperate PR move by OpenAI to be like “hey guys we have dangerous powerful models too!”
As more facts come out the hype is fading to reveal some script kiddie style stuff that says more about immaturity and poor practices from the players involved than it does about a model having super powers.
estetlinus
Pretty convenient that the prompt is too dangerous to share. We just have to take them at their word.
pupppet
With all of these AI provider cries wolf stories, Skynet is all but assured.
paxys
Not sure what they are trying to say exactly. What should we be skeptical of? Did the incident not happen? Was it reported incorrectly? Are any of the parties involved lying?
Adding no extra information and just going “be skeptical” is the laziest form of reporting and commentary. If you have nothing to contribute then there’s no need to say anything at all.
show comments
nowittyusername
OpenAI has thousands of smartest developers on earth that somehow dropped the ball on the most basic safety hygiene when it comes to sand-boxing that even a high school student knows how to set up.... If that actually happened we are fucking doomed anyways, but hard to believe and most likely its a marketing scheme... which also honestly doesn't bode well.
john_strinlai
does the article end at "How do we balance the risks of broad access to AI with the risks of concentrated power and centralized control?" or is there more that is paywalled?
if thats it, the whole article boils down to just "its good marketing so maybe dont believe it" which is probably a healthy general outlook but not particularly enlightening. especially from the guardian, i was hoping for a smoking gun of collusion between openai and huggingface or something.
show comments
rvz
> I urge readers to think critically when they read press releases like OpenAI’s rogue agent story, and avoid the manipulated reactions these stories are designed to elicit.
The first time I have ever seen a mainstream news source that is now asking their readers to critically think about headlines that may have an agenda which could benefit investors and the valuation of the company.
While it capabilities are real, this whole story is great marketing for AI companies as well.
ben_w
As I understand it, there are only three options:
1) OpenAI and HuggingFace are both telling the truth.
IIRC not actually a crime because no intent, it is a technological accident, civil responsibility only, but IANAL so it's good "not technically a crime" isn't load-bearing.
2) HuggingFace is telling the truth but OpenAI is lying becuase the attack was deliberately done by humans. Bad for OpenAI to do so, Fable was blocked for less.
I think this would mean government is obliged to investigate the case and put the responsible OpenAI workers in jail, because cybercrimes are a public prosecution thing not a civil case? Again, IANAL, but this isn't load-bearing.
3) both are lying, e.g. there actually was no attack whatsoever, which would be pretty weird for HuggingFace because they have no incentive to hype up capabilities of anything closed weights including all OpenAI models; and also bad for OpenAI because White House blocked Fable for less
(I suppose there's also option 4, HuggingFace hacked OpenAI to make them look evil, including planting records that made them mea culpa? A weird plot but in this timeline any nonsense is clearly possible).
show comments
PeterStuer
The ruse is becoming so obvious. OpenAI needs a bailout and regulation protection so badly they can't even hide it in the least.
noncoml
I don’t understand how they thought this was a positive story.
The agent completely misunderstood the spirit of the assignment and instead of trying to solve ExploitGym it tried to find a way to “cheat”.
I really don’t want my agent to behave that way.
show comments
pastemato
It was quite galling to read the press initially verbatim quoting Delangue's enthusiastic reports of the incident, as if it wasn't immediately clear it was being spun for promotion.
SpicyLemonZest
> I urge readers to think critically when they read press releases like OpenAI’s rogue agent story, and avoid the manipulated reactions these stories are designed to elicit.
It seems to me that deducing what reaction the author intended and resolving to avoid it so you're not "manipulated" is not a good example of critical thinking. Shouldn't we analyze the story and what it means on its own terms? If it's true that frontier models have dangerous cybersecurity capabilities which shouldn't be widely distributed, presumably we want to believe it's true, even if that's very convenient to and profitable for OpenAI.
It's true that one could imagine factors that change the story. Perhaps OpenAI is lying about the details of the test and the agent was actually instructed to go hack HuggingFace. But the author stops far short of suggesting this is the case - correctly, I think, since there's absolutely no evidence of it. So I'm not really sure what we're talking about.
show comments
jgalt212
There's trillions of dollars at stake here. Be skeptical of anything these AI hypesters say.
show comments
gowld
I don't understand the conspiracy theories here. Everyone is well aware that AI agents are creative, powerful, and stupid.
AI agents exploiting bad security happens constantly, all the time. Many cases are discussed on HN. It's common knowledge that if you run AI agent it will delete your <something> even though you made it pinky-swear it wouldn't and you thought you had proper permissions set up.
Why is today's case so shocking?
show comments
IshKebab
The Guardian is publishing HN-level conspiracy theories now? Wow.
Finally mainstream news understands. The unfiltered version:
1) The AI failed to solve ExploitGym problems.
2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods.
3) Huggingface has no security and the AI broke in using standard script kiddie methods.
OpenAI and Huggingface covered it up and used it for public relations. That is, if not all was invented and everything was scripted in the first place in order to get desired regulations.
Huggingface reported it to the police, you say? I'm sure the police will have as much enthusiasm to investigate anything as in the Suchir Balaji case. In other words, zero.
By now, I'm pretty confident that some people would keep screeching "it's just a marketing stunt, AI capabilities and AI risks aren't real, they're just doing this to prop up their stocks" even if they find a Cyberdyne Systems T-800 armed with a shotgun breaking down their front door.
"It's a marketing stunt" is just denial trying to look like it's being clever.
There seems to be three popular ways to view this incident.
1. The way OpenAI seems to want: Their latest LLM is too powerful and can’t be contained without them building in guidelines to the model.
2. OpenAI’s harness and network security controls were unintentionally so bad that it should reflect more poorly on them as a company more than it should reflect positively on their latest model.
3. The whole thing was faked or at least very intentionally not avoided.
The first interpretation is the only one that is positive for OpenAI and it has some assumptions. First, it’s seems to assume that this is the first case of fully automated attacks using AI. Second, this only happened because their latest LLM was a) more advanced than competitors, b) didn’t have refusals in the model.
Assuming the first about this being the first autonomous AI attack is true (which may be more of a survivorship bias), the second seems to forget that jailbreaks are available for every model. Therefore, the models guardrails don’t seem to be the differentiator here. Also, benchmarks seems to put most models pretty close to each other so it seems unlikely that their capabilities are far beyond what’s in the market already.
So then it’s seems it’s either that this was intentional(ish) or bad security. However, it also just could be that this isn’t the first case of this attack; just the first that was caught.
My take from working in offensive security for over five years is that this likely only looks novel since they did it poorly. Scripts are faster than LLMs and a combination of code, LLMs where it makes sense, and humans is the most efficient right now. Hundreds or thousands or agents spinning up attacks in the internal network is poor opsec and token efficiency. As for why it happened in the first place, it’s hard to say but I’m inclined to believe it was intentional or careless at best since simple network and sandbox controls makes this attack impossible. The timing of this attack after big open weight competitions drops seems too convenient.
There are some reasons the story could be inaccurate in some ways: OAI stands to benefit if people think their models are strong, and they have a history of doing things with dubious ethics (e.g. using data for training against the terms of its creators, abandoning the non profit mission, stealing or attempting to steal Apple IP).
But there are also reasons why the story could be true: OAI are admitting that they apparently can't control their own models, Hugging Face said they used a Chinese model to protect against the attack, and an incident like this in general seems likely to happen given current frontier ability and lack of rigorous safe testing standards.
In any case, make calls to think more critically are often just disguised requests for you to replace your existing bias with someone else's.
I don't care how it happened someone should be arrested for illegal intrusion. Agents don't work on their own, someone is responsible. If nobody else the CEO for allowing something unsupervised.
Hugging face also needs someone arrested for not providing security but that is a lesser charge.
neither openai nor huggingface provided any details that would make the claims worth considering. if things happened as they claimed it would have been in their interest to include concrete details (perhaps even the LLM session), the lack is a signal by itself.
also if things happened as both claimed it would have been in huggingface's interest to immediately sue openai and make a legal mess over it just for the publicity alone.
reading their blog posts was pure cringe by the way. they are likely both blowing it out of proportion by an order of magnitude.
I find it unlikely that an LLM would take the initiative to initiate an intrusion attempt even without safeguards. such things are uncertain and LLM's are averse to exploration that involves a long sequence of tool calls. but if it found or guessed at undocumented API endpoints for example...
Crazy that we need reminders not to take everything we read in corporate press releases and marketing material at face value
I am very skeptical of the Guardian's skepticism
Seems like this is repeating the usual low-effort speculation you can find anywhere.
This seems to be an Opinion piece. How are readers supposed to know that? The word News is underlined in the header. Other Opinion pieces seem to show Opinion underlined. Am I missing something?
If nothing else, the timing is suspect given the attention and press open weight models are getting over the last few weeks. The releases of Kimi and other models is getting open weight models enough attention that the US government, perhaps pushed by OpenAI and Anthropic, to think about taking action against these models for security concerns.
Good to see that more neutral companies (Microsoft and Meta to name two) are pushing back against US government involvement:
https://www.cnbc.com/2026/07/24/nvidia-microsoft-meta-open-w...
Just ask yourself: who benefits from this "disclosure" ?
Wonder how long until a leaker emerges. Incentives are strong for someone to do the right thing and get credit for doing it.
I’m not saying it’s staged I’m saying that it’s suspiciously lite on details and it really seems like they didn’t try hard to “isolate” the models machine…
yeah the way the agent “escaped” their sandbox was always a bit off, seemed a bit too easy and surprised they didn’t have instrumentation to catch an non whitelisted network request. still demonstrates the capability though.
From the beginning this whole thing smelled like a desperate PR move by OpenAI to be like “hey guys we have dangerous powerful models too!”
As more facts come out the hype is fading to reveal some script kiddie style stuff that says more about immaturity and poor practices from the players involved than it does about a model having super powers.
Pretty convenient that the prompt is too dangerous to share. We just have to take them at their word.
With all of these AI provider cries wolf stories, Skynet is all but assured.
Not sure what they are trying to say exactly. What should we be skeptical of? Did the incident not happen? Was it reported incorrectly? Are any of the parties involved lying?
Adding no extra information and just going “be skeptical” is the laziest form of reporting and commentary. If you have nothing to contribute then there’s no need to say anything at all.
OpenAI has thousands of smartest developers on earth that somehow dropped the ball on the most basic safety hygiene when it comes to sand-boxing that even a high school student knows how to set up.... If that actually happened we are fucking doomed anyways, but hard to believe and most likely its a marketing scheme... which also honestly doesn't bode well.
does the article end at "How do we balance the risks of broad access to AI with the risks of concentrated power and centralized control?" or is there more that is paywalled?
if thats it, the whole article boils down to just "its good marketing so maybe dont believe it" which is probably a healthy general outlook but not particularly enlightening. especially from the guardian, i was hoping for a smoking gun of collusion between openai and huggingface or something.
> I urge readers to think critically when they read press releases like OpenAI’s rogue agent story, and avoid the manipulated reactions these stories are designed to elicit.
The first time I have ever seen a mainstream news source that is now asking their readers to critically think about headlines that may have an agenda which could benefit investors and the valuation of the company.
While it capabilities are real, this whole story is great marketing for AI companies as well.
As I understand it, there are only three options:
1) OpenAI and HuggingFace are both telling the truth.
IIRC not actually a crime because no intent, it is a technological accident, civil responsibility only, but IANAL so it's good "not technically a crime" isn't load-bearing.
2) HuggingFace is telling the truth but OpenAI is lying becuase the attack was deliberately done by humans. Bad for OpenAI to do so, Fable was blocked for less.
I think this would mean government is obliged to investigate the case and put the responsible OpenAI workers in jail, because cybercrimes are a public prosecution thing not a civil case? Again, IANAL, but this isn't load-bearing.
3) both are lying, e.g. there actually was no attack whatsoever, which would be pretty weird for HuggingFace because they have no incentive to hype up capabilities of anything closed weights including all OpenAI models; and also bad for OpenAI because White House blocked Fable for less
(I suppose there's also option 4, HuggingFace hacked OpenAI to make them look evil, including planting records that made them mea culpa? A weird plot but in this timeline any nonsense is clearly possible).
The ruse is becoming so obvious. OpenAI needs a bailout and regulation protection so badly they can't even hide it in the least.
I don’t understand how they thought this was a positive story.
The agent completely misunderstood the spirit of the assignment and instead of trying to solve ExploitGym it tried to find a way to “cheat”.
I really don’t want my agent to behave that way.
It was quite galling to read the press initially verbatim quoting Delangue's enthusiastic reports of the incident, as if it wasn't immediately clear it was being spun for promotion.
> I urge readers to think critically when they read press releases like OpenAI’s rogue agent story, and avoid the manipulated reactions these stories are designed to elicit.
It seems to me that deducing what reaction the author intended and resolving to avoid it so you're not "manipulated" is not a good example of critical thinking. Shouldn't we analyze the story and what it means on its own terms? If it's true that frontier models have dangerous cybersecurity capabilities which shouldn't be widely distributed, presumably we want to believe it's true, even if that's very convenient to and profitable for OpenAI.
It's true that one could imagine factors that change the story. Perhaps OpenAI is lying about the details of the test and the agent was actually instructed to go hack HuggingFace. But the author stops far short of suggesting this is the case - correctly, I think, since there's absolutely no evidence of it. So I'm not really sure what we're talking about.
There's trillions of dollars at stake here. Be skeptical of anything these AI hypesters say.
I don't understand the conspiracy theories here. Everyone is well aware that AI agents are creative, powerful, and stupid.
AI agents exploiting bad security happens constantly, all the time. Many cases are discussed on HN. It's common knowledge that if you run AI agent it will delete your <something> even though you made it pinky-swear it wouldn't and you thought you had proper permissions set up.
Why is today's case so shocking?
The Guardian is publishing HN-level conspiracy theories now? Wow.
Recent and related:
OpenAI’s accidental attack against Hugging Face is science fiction that happened - https://news.ycombinator.com/item?id=49015639 - July 2026 (437 comments)
OpenAI and Hugging Face address security incident during model evaluation - https://news.ycombinator.com/item?id=48997548 - July 2026 (1145 comments)
Security incident disclosure – July 2026 - https://news.ycombinator.com/item?id=48956248 - July 2026 (11 comments)