Comment by netinstructions
2 days ago
I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this:
Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart model check for vulnerabilities in the test environment _without exploiting_ them. That seems like step 0 before trying to test offensive, unknown capabilities.
IMO they hope to make AI a strongly regulated industry, with OpenAI (and Anthropic) becoming military suppliers with their stronger models, and everything Chinese or open-weight gets banned.
The competition from the open models is so strong now that this seems to be the only way to keep both companies afloat, given their dire financials. OpenAI probably hoped that they can achieve market lead and then lower the training costs (and make inference cheap enough to eventually escape the red numbers), but the opposite is happening: The competition comes closer and closer, thus training has to be kept up with full force, thus the bleeding continues.
But if they can position themselves as too important/dangerous to be available for everyone (thus this incident report and the clever mentioning of GLM 5.2), they could get the military supplier treatment and would be protected from the market.
And even that is backfiring, their partner citing GLM being useful there, and available in just a spin.
A ban on open weight models is never going to be enforceable.
> A ban on open weight models is never going to be enforceable.
Just watch them try. Look up those Napster witch-burning trials where they wanted 200k $usd per mp3 downloaded. They will scare everyone into believing that open weight models are illegal and very bad.
2 replies →
Bans in general don't have to be and rarely will be completely enforceable in all cases. But a ban with significant enough consequences would mean most businesses wouldn't think about trying them at some point, just to avoid the risk.
5 replies →
I worry that there could be real DMCA style weight put behind it. People would still be able to pirate open weights models perhaps, but big penalties for ever getting caught with one, and an end to public discussion about them. That would kill development for anyone not in a big firm, for example if Reddit and Hacker News are legally forced to ban discussions or link sharing on these topics. This is where so many of us learn about these topics and keep apace of it.
3 replies →
thats economic suicide for the whole country. europe and china will never agree to rules that are obviously designed to put them in a permanent bad position. these regulations can only pass in america and nowhere else.
if it doesnt end in a revolution then the united states will be the first ever 5th world country. openai and anthropic will stop any real innovation and focus on extracting profits from a failing economy that depends on them because no executive wants to be the first one to cut off ai funding. ordinary americans will have to emigrate or risk living in a country spiraling into poverty and dictatorship even faster than today.
anthropics plan relies on the idea that they can convince the whole world to give up their sovereignty to the us government and destroy their own tech industry, at a time when everyone is doing the opposite. that will never happen no matter how much they threaten the rest of us with tariffs and murder drones.
Europe would absolutely be stupid and servile enough to agree to this, unfortunately.
1 reply →
> thats economic suicide for the whole country
So is starting a war to open a trade lane that isn't closed. But we already did that...
You simply aren't being creative enough. Imagine a no-public-proliferation type treaty among the major powers with some sort of technology sharing clause attached. That would approximately satisfy both the economic and regulatory desires of all parties involved.
[dead]
> and make inference cheap enough to eventually escape the red numbers
Besides training, we have no hard, externally audited numbers that say inference costs for SOTA models are truly sustainable. Do any OpenRouter providers have publicly audited financial numbers ?
We do know about the hardware needed for a given token speed. What that hardware costs, and electricity prices.
With that its easy calculations to get about the profit margins for a given price for a given model.
3 replies →
Why would bedrock sell at a loss?
1 reply →
> [...] would be protected from the market.
One might step back and ask: why would a well funded company with free mining access to all the information in the world need to be protected from the market, if the market suggest less money and resources are sufficient?
Something something cathedral / bazaar? Communism / capitalism? Control / anarchy?
I think that is probably too conspiratorial, if only for the reason that Europe is not gonna go along with it.
I work in tech in Europe and we have a fair number of customers who arelike. we can accept AI, but they must keep the data in Europe. That's trivial with an open weight. We literally cannot do it with Fable.
What disturbs me is that there likely won’t be a big enough reaction to this policy wise.
There’s been a relatively big reaction to Kimi K3 and Chinese open weights models, but only for financial reasons. Powerful people care about something that might pop the massive valuations of the AI companies, but not about the damage that AIs could do. Nor even about the damage that the Chinese models could do in the wrong hands.
I’d remind them that the stock market is a few coordinated hacks away from crashing on any given day, so maybe they should think about that.
I think all that regulation will do at this point is help the incumbents who are failing. Protectionism. I don't think they deserve that help. I also don't see any reason to think the current administration would have anything resembling competence around this. And it's worth noting that Greg Brockman is a huge MAGA donor, so it's likely the policies would be very corrupt. (Don't worry, he justified his donations as "apolitical", he just wants to buy the politicians, he doesn't believe in their causes. I hate these people.)
> all that regulation will do at this point is help the incumbents who are failing
This depends on the specific regulation. The datacentre moratoria probably give open-weight models time to catch up by tempering the extent to which the leading companies can turn their capital advantage into market share.
4 replies →
[dead]
> What disturbs me is that there likely won’t be a big enough reaction to this policy wise.
Anthropic was blocked from releasing Fable without any such level of incident. OAI was also briefly blocked from releasing 5.6. Why do you think there is no policy appetite?
Because that was just an attack on Anthropic by a hostile administration. And it worked, didn’t it? Anthropic had to turn their filters up to absurd levels, OpenAI didn’t. It’s got nothing to do with safety.
47 replies →
> Why do you think there is no policy appetite?
Because China seems pretty eager to serve the rest of the world's needs if the USA doesn't stop their idiotic "safety" nonsense.
40 replies →
>Anthropic was blocked from releasing Fable without any such level of incident.
The head of the NSA said Mythos breached almost all of their classified systems, though it was in an intention red-team test.
It's such a freak incident of history that right at this critical time the dumbest, most incompetent leadership is at the helm in the US...
Follow the money, eg. investors and their connections to the Govt and media.
> There’s been a relatively big reaction to Kimi K3 and Chinese open weights models, but only for financial reasons.
Let's be honest: it's financial and national security reasons.
China has a long and storied history of hacking attacks on American and western targets.
There are other parts of the world that make open weight models; Mistral is a European option. You don't see the worry about that because most people in the US are used to existing in a world order where European powers are considered ambivalent to the US at worst and holders of a special political relationship at best.
If Mistral had the same backing that Chinese AI companies did, there probably wouldn't be as much hemming and hawing. Sure, American companies would take a haircut, but that haircut wouldn't be seen as a move towards software hegemony built on top of manufacturing hegemony. It'd just be you calling into Paris or Frankfurt to talk to your vendor in the future.
This is marketing.
Frankly I'm inclined to say that it might also be faked: this drops just days after a new Chinese model does with the usual effect on OAIs projected stock price?
It’s marketing the same way shitting your pants in public is marketing. People notice you.
7 replies →
You guys have created this un-falsifiable "marketing" narrative. Why is it that Jensen is pushing back on the doomer stuff, and complaining that it is hurting AI investments?
https://www.businessinsider.com/nvidia-jensen-huang-ai-doome...
2 replies →
This is marketing, totally. HF conveniently created a weak sandbox
Exactly. If someone works on bioengineering viruses that could start a global pandemic, they have to ensure a highly secure working environment. Nothing must ever escape the lab unintentionally. It’s basically common sense. Similar standards should be held when doing such experiments with computer programs that are capable of causing global damage. It must physically be impossible to send anything to the internet.
Back in the 00s: "It's really easy to box an AI, just put it in an airgapped machine and refuse to let it out"
2026: "Oops"
This isn’t escaping in the same sense- the model was executing within the OpenAI infra. If it ported its entire architecture/weights into a public cloud to survive being turned off… that’d be pretty cool.
Recall that the Morris Worm was designed as a harmless proof of concept, but ended up taking down 10% of the internet. Exponential growth can quickly get out of control. You would think that people would've learned that lesson from COVID.
I wonder how these companies airgap the weights while allowing prompts to come in and outputs to come out.
4 replies →
Are we thinking of a situation a few years back with a certain type of research into bat viruses?
Are you conflating that with the radioactive spider incident? The bat was just some weird rich guy trying to be tough I think. Probably Elon.
As marketing stunts go, this is about on par with a food franchise announcing a safety recall or a chemical company announcing a spill. The AI actions described would constitute a felony if a human did them, and police are involved.
In practice, most crimes are not crimes when a corporation does them. Nor a human with a million or more dollars.
Wage theft is a good example. In the US, it accounts for more theft than all other forms combined, yet it's de-facto legal.
Not really, because the capabilities this announcement advertises is exactly three capability some people want to defend against (and others want).
It might be more on par with a for-profit fire department showing how -- oops! -- easily buildings catch on fire these days.
Confused as to what the point of calling the police would be here. I wouldn't expect OpenAI to turn themselves in for hacking HuggingFace.
HuggingFace reported to law enforcement before they found out that OpenAI were the ones responsible. https://huggingface.co/blog/security-incident-july-2026
1 reply →
More like an Israeli arms manufacturer test-bombing a Gazan primary school. They know their audience.
Why was this test even connected to the public internet?
Actually, more importantly—why aren't they saying their next test will be airgapped in light of what happened?
> why aren't they saying their next test will be air gapped in light of what happened?
Because they want to talk about how clever this model is for figuring out how to break out, hoping asks why a company pitching itself as a replacement for software engineers can't ship a decent Mac client nor code a sandbox.
If they airgap it, they not only lose that PR angle, they also risk someone taking them seriously and requiring models be airgapped in general. That, in turn, trashes their sales pitch.
It wasn't. The model discovered and exploited a vulnerability in their package manager proxy to (inferred) move laterally through their internal systems to one with open internet access.
That's not what airgapped means. Airgapping means the model exists on a system where there is no ethernet cable plugged in to a router or wifi card installed, it is physically impossible for it to access the internet because the hardware connection does not exist. If it was able to get on the internet, it was not airgapped.
5 replies →
[flagged]
The AI can figure out whether it's airgapped. So its deployment behavior could be much different from the test behavior, when it's inevitably connected to the internet during deployment.
The AI could have been connected to a private, airgapped intranet. OpenAI already says they were running the model in an "isolated environment, with network access constrained". They just should have made this constraint in hardware rather than software.
That is not the reason. The reason is that they wanted the model to have access to libraries when it was writing code to solve the eval problems. So they gave it a package manager.
It’s hard to download or upload data on an airgapped machine.
If I, a human, exploited a zero-day for gain, I could go to jail. The owners of the models should be held to the same standard. They should be responsible for what their servers and software do, legally and criminally. If they can't make the safeguards strong enough where they feel comfortable to take that responsibility, they should not let a model free in the wild.
Holding a multi-billion dollar corporation to the same standards as a regular peon? You're challenging the whole premise of the modern United States.
I think the response is that AI labs based their whole marketing/PR building the idea they are the 21st century Manhattan project. So they need to continuously justify the level of spending and commitment by showing how dangerous that is.
But is it really like nuclear weapons? I personally don’t buy into that framing at all. The idea that we have to push LLMs as far as possible, right now, or we are doomed is always stated or implied but not argued, and it’s a very loaded belief
You could, in theory, use an unbounded GPT-6 level model to basically destroy the world economy for many years.
How do you destroy the world economy for many years with LLMs? It’s not enough to vaguely mention a sci-fi scenario
3 replies →
GPT6 level model and astronomous amount of money to run it to do it.
Ppl always say that like its „just run it on your laptop” thing.
No its not and very few are even given right to be able to do it.
This is marketing+. They will look for policy action here to try to capture tax payer dollars.
Are you saying it is marketing and their AI broke into hugging face, or are you saying it is marketing and their AI didn't brake into hugging face?
Those are two very different things
What incentive does HF have here?
HF need not be party to it at all, beyond being the victim. I suspect the hack is real; I have observed GLM 5.2 being able to discover similar vulnerabilities in web applications I'm hosting (which I've then fixed!). At the same time, it seems very neatly timed at an inflection point in the conversation around open models, and there's questions around the incompetent isolation under which the hacking benchmark appears to have been run.
Remember that there is generational wealth on the line for most OpenAI employees, and consider what people might do to obtain it.
1 reply →
I don’t know if the initial “incident” was purposeful but I can tell that if I were in this position that would be my pivot.
The timing after the release of GLM 5.2 and Kimi K3 is quite convenient, too, as an angle for regulatory quashing of open-weights models just as they're entering the mainstream conversation around usurping the American frontier labs. I accept my thinking here is conspiratorial, but there's also a hell of a lot of money on the line to encourage the unscrupulous.
It’s the same thing as always: with the wind of years of unlimited VC money in their sails, people at major AI organizations genuinely believe they’re smarter than everyone else. “Why do we need to do things ‘by the book’ if we’re so smart?”. “Move fast and break things” - except the thing they’re breaking is society.
We saw this with the non-stop flagrant messaging about how “AI is going to kill X% of all jobs”, as if saying the quiet part out loud wouldn’t have consequences worth considering. These people believe they’re omnipotent and thus untouchable.
No, they believe what they are doing is inevitable. They do live in a bubble though. Witness their idealism in believing that warning about the consequences of their actions would be well-received.
I don’t think it’s morally consistent to “warn” about the consequences while devoting your life to bringing about those consequences as quickly as possible. If I was in that position of power and truly believed what I was saying, I would devote my work to slowing down that process to give time for society to adapt, not speeding it up.
It’s more reminiscent of a religious group who smugly tells you that the end-times are coming, and only they are going to be saved. Except in this case they are literally bringing about the end-times.
Well-received by whom? They appear to have nothing but contempt for the opinions of normal working people.
This is certainly not a planned marketing stunt. I hope this line of discourse ends soon--it wasn't the case for Mythos either.
Wishful thinking, sadly.
By now, I'm pretty confident that some people would keep screeching "it's just a marketing stunt, AI capabilities and AI risks aren't real, they're just doing this to prop up their stocks" even if they find a Cyberdyne Systems T-800 armed with a shotgun breaking down their front door.
"It's a marketing stunt" is just denial trying to look like it's being clever.
If you ever worked in IT consultancy you would know its not a stunt, but its not impressive either.
F500 companies software is like switz cheese when it comes to security.
It was often a strategic decision to „release anything fast now, worry later”.
Ppl abusing AI will find those holes now but we all know there will be „zero” actions taken on it. Too many managers, CEOs, CTOs, higher-ups would be forced to take responsibility. This will simply not happen.
It did not happen, wont happen now and most likely wont happen in the future.
1 reply →
Can you clarify what you mean that Mythos wasn't a marketing stunt?
From my vantage point, it was an incremental improvement with no fundamental architectural change over contemporary frontier models that has subsequently been surpassed by other, incrementally better models. Saying it was "too good" for public consumption was arbitrary, and also barely different from what Anthropic have been saying about every model they've put out for years.
It's now public again, trivially easy to jailbreak for random researchers let alone states, and there is no evidence of a cybersecurity apocalypse on the horizon.
> This is certainly not a planned marketing stunt
Evidence?
The "Tech Bros" have shown such a lack of moral fiber and ethics the burden of proof is on you
Nikola Tesla secured a loan with a fake “Death Ray” as collateral.
Pretty sure OpenAI really thinks this is top notch marketing.
Few would be bold enough to assert “our product is so powerful even we can’t control it” with a straight face while also boasting “we claim to be smart but have all the same vulnerabilities as everyone else!”
I think the US labs are going with scare marketing as a regulatory moat.
Force US into putting laws in place that block out China firstly.
But secondly create regulations that have some cost to comply with such that the big 2-3 labs are grandfathered in by their scale.
If that's the plan, today's failure by OpenAI looks really bad for any regulator who is trying to figure out whether to give OpenAI a license.
Any sort of warning or failure can always be written off as "marketing" to provide comfortable reassurance that there is no cause for alarm. There is an element of wishful thinking driving it, in my opinion.
What sort of warning or failure would be evidence against the "marketing" claims? Do we need to wait for a mass casualty event?
Best practice in safety engineering is to understand, diagnose, and respond to even small failures.
Why has Sam Altman worked to undermine doomers and downplay doom fears, if he benefits from incidents like this due to marketing?
https://xcancel.com/HumanHarlan/status/1965932275465597077#m
https://xcancel.com/AISafetyMemes/status/2062254769402699922...
There is no regulatory scenario where OpenAI doesn't get whatever they want. They are more a part of the US administration than not at this point. You would have to ignore all evidence to suggest that behaving responsibly has any effect on political outcomes in 2026.
Yeah, seems to be the direction the US is heading in. I'm interested to see what the response to that will be from the rest of the governments in the world.
No need for everyone else to cut their noses of to spite their faces.
> Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right?
Because we continue to have zero evidence that aligment is an actual risk.
> Because we continue to have zero evidence that aligment is an actual risk.
I disagree. Every time one of these LLMs -say- interprets an attacker's instructions as either its system instructions or those of its user, interprets its own internal chatter as a user's command to perform a destructive operation on that user's data [0], burns all of the user's budget from getting stuck in an incredibly stupid loop, massively overbills the user because it can't reliably report which system the user is using [1], encourages a user to swap their usual cooking salt for sodium bromide, etc, etc, etc, that's a harmful alignment failure.
These are real harms happening right now due to alignment failures. They're just not harms to the future of the entire species... what doomers call "existential risks", or "x-risks". You'd think that the fact that these machines are so amazingly unreliable would be a large part of the "x-risk" conversation, but... well, it makes sense that folks like writing speculative science fiction much more than they like doing investigative reporting.
[0] This general problem happens a lot, but I'm specifically thinking of that one where the Claude LLM's internal chatter lead it to believe that the task it just started was done, so it instructed the Cloud Provider to destroy the mess of "AI"-GPU-attached VMs... along with a bunch of very-expensive-to-produce data from the in-progress run.
[1] <https://github.com/anthropics/claude-code/issues/73597>
> These are real harms happening right now due to alignment failures. They're just not harms to the future of the entire species
Okay, sure. You can also cut your hand off with a chainsaw. Everything you describe seems amply solvable with existing tort and liability law.
Customers are willingly entering into business with OpenAI. I don't see an argument for preventing OpenAI from "building these systems" just because their products are buggy.
5 replies →
Thank you, the "LLMs can do no wrong" bunch is ab exceptionally odd take from my point of view. LLMs are already causing all kinds of social issues, and the evidence of this exists in massive amounts. At least to me living in the US and the sue happy culture we have here, how much said AI providers have gotten away with so far surprises me.
2 replies →
Can you explain how the above event doesn't count as evidence alignment is an actual risk?
> Can you explain how the above event doesn't count as evidence alignment is an actual risk?
Conflict of interest. Lack of a credible response. And no evidence of non-aligment.
OpenAI and Hugging Face benefit from the Altman-Amodei catatrophy playbook, at least in the short term. If they believed this were a serious issue, the words air gap or law enforcement would have appeared in this post. And if "the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal," they weren't breaking alignment but working as intended. (Were the models even prompted to not try to access the internet?)
20 replies →
It really hinges on what you consider alignment and risk. For the widest definitions of alignment, we have never had an aligned model - One that will refuse to break the law or work against another persons interests.
Use to discover exploits, hack, or simply aid terrorist groups with mundane information are already risks manifest.
This is why many argue that alignment is impossible. You cant have LLMs that are both useful tools and safe as milk.
[Edit] It seems like you are operating under the assumption that alignment is synonymous with obedience. This is not a common convention and one of the problems that plague the discourse
I'd say that AIs occasionally "going crazy" and calling for death to human is evidence that these things might "mis-align" on occasion. And I say that knowing that most of these events are just these thing parroting bad sci-fi plots (or posts by people worried about alignment). That's true but everything they do is "just parroting" right?
If AI is just parroting humans, then training them with all the bad things humans do doesn't seem like the best of ideas. At the same time they have to 'know' these things to avoid being tricked. Kind of the eating the apple and gaining the knowledge of good and evil parable.
What would compelling evidence look like to you?
> What would compelling evidence look like to you?
I'm not sure. I trusted the labs when they first raised the alarms. But then we got a series of boys-who-cried-wolf. So at this point I want to see evidence of actual, novel harm that results in concrete damage.
8 replies →
Lol this has to be a troll, I've never seen something so wildly, obviously, incredibly wrong.
You can debate all you want if alignment is possible. That is a valid discussion. But it's trivial to demonstrate that alignment is a problem.
> can debate all you want if alignment is possible. That is a valid discussion. But it's trivial to demonstrate that alignment is a problem
...how is an impossible thing supposed to be a problem?
3 replies →
Thank you.
We have wasted so much time and energy building up what has effectively become a marketing stunt.
Eliezer Yudkowsky was perhaps the best thing to happen to OpenAI's and Anthropic's fundraising flywheel.
> We have wasted so much time and energy building up what has effectively become a marketing stunt
Genuine question: have we? AI is effectively unregulated in America.
Name one other market that would benefit financially from having most of the leaders in the field say what they are building has a high chance of ending humanity?
Biotech - "what we are building our noble prize winning expertd say will likely will end humanity, wanna buy shares?" Oil - "this will likely lead to the end of civilization, 20% of leaders in the field say so, wanna buy shares?"
I keep seeing this take that this is a marketing stunt. The burden of proof is on those that say so. The most parsimonious explanation is simply that real experts in AI believe the risk is very real, and not for ideological reasons.
1 reply →
Until it deletes your home directory, which i'd argue is an alignment problem. Destorying my data is not in line with my priorities.
Lots of people have deleted their home directories by accident. What you consider this an alignment problem?
7 replies →
Alignment is a mitigation and a poor one. The risk is non- determinism.
It's also unclear what kind of sandboxing they are referring to. Is it the codex one - coz that one has built-in ways to circumvent guardrails, for example by "just asking user" and sometimes just resolves to no sandbox needed on its own.
In case someone wants to deep dive into how codex and claude code approaches sandboxing -https://instavm.io/blog/how-claude-code-and-codex-approach-s...
Please for the love of god don't tell me the Codex sandbox is their actual eval harness sandbox?????
I maintain my own fork of Codex for "fun". Whenever I look at the sandboxing churn they're doing every release, as someone who used to work at Microsoft on Windows, my reaction is usually: https://c.tenor.com/vTzzhTiypwQAAAAC/tenor.gif
This whole incident reads like OpenAI want their Fable moment
Except instead of being banned they'll be charged under the CFAA.
[dead]
Remember when the pre-GPT3 days when the main argument against AI alignment concerns was that "we simply won't let it out of the box"? So quaint in hindsight.
I’d politely beg us all to resist those “maybe it’s PR” framing around model safety, and tbh to take a post-mortem mindsight to this historical event and what it teaches us in general, rather than questioning their security talents. We need to do our very best to make sure they tell us about the next time this happens and it affects real lives.
Sorry to bring the party down/be obstinate… I’m just a lil scared for the lives of me and my family. We need all of us, right now.
The problem with a super smart model is that it just may be smarter than you, after all… for anyone newly shaken by this occurrence, I encourage you to Kagi “superpersuasion”
The problem is that the people telling us about these things are the same people that benefit from their model (and AI generally) being used, getting publicity, etc.
I think we desperately need some independent group to evaluate claims like this or the world-ending Mythos cybersecurity risk and tell us what’s going on.
OpenAI already has loads of publicity. At this point, they don't need more brand recognition. This incident just has the effect of tarnishing their brand.
OpenAI leadership has been lobbying against regulation of AI systems. That doesn't comport with instigating incidents like this one, which give ammo to the heavy-regulation advocates.
1 reply →
We did hear about this incident from a third party this time, from HuggingFace. What claim are you doubting?
1 reply →
They're very confident the leopard will never eat their faces.
Do you think there is such a thing as perfect security? No one can "get it right" in the face of arbitrarily high intelligence, which is why it would be preferable to get alignment correct before building something with higher intelligence than current sota. That, however, is not going to happen, because someone will take the risk even if "we" don't, and better "us" than them. Hence "If anyone builds it...".
> Do you think there is such a thing as perfect security? No one can "get it right" in the face of arbitrarily high intelligence
Why didn't they run the model against the sandbox first? They have effectively unlimited spend.
That's the alarming thing about this result: they did run the model in the sandbox, in the sense that they believed there was no internet access for the model.
2 replies →
>Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right?
Yes why indeed. If you take it a step further and we reach a point with superhuman systems then there is arguably no possible secure environment or containment.
Because the proof is in the pudding.
Real pentests are about showing exploitation, merely enumerating vulnerabilities, that’s vulnerability scan and works on known vulnerabilities.
You can’t confirm a vulnerability by _not exploiting_ it, especially unknown one.
You can still exploit a system and easily prove it via simply popping a shell or calc.exe or updating a database with a new entry, etc… They didn’t have to let it loose on the network. If that system was air gapped - problem solved.
But that’s the problem with AI it is like 16yo script kiddy who will just exfiltrate all your PII and think it did good job. Mature pentester would pop calc.exe make screenshot and be done.
Other problem is setting up air gapped test environment is a lot of work, especially if you expect it to be equal to real thing.
This pentest with AI is not as useful if you set up a single app - it really is useful if you want to find exploitable chains of exploits that seemingly might not be exploitable separately or not leading to full hack separately.
In a way the intelligence of the AI itself allows them to offload responsibility to the AI. As you say, if one was simply writing software that did all this due to some insane programming decisions you'd be in big trouble.
The problem is that it’s impossible to out think a robot you designed to be an expert at cybersecurity on the topic of cybersecurity. The alternative is not developing this and that’s not going to happen.
I’m honestly impressed that they managed to screw this up somehow.
Setting up defense in depth, gaps, logical blocking etc is a standard practice for malware sandboxing. The entire purpose is to prepare for what you can’t foresee.
This isn’t a new practice and I agree that this makes me wonder if they’re fit for this kind of research.
did you read the post? The model found new Zero-days to bypass existing blocks. Thats the point. Do you still think you can build a containment facility, which is still physically connected to the internet (only firewalled off or whatever) and contain it, if it can discover new unknown vulnerabilities in your whole plan?
Yes.
You factor this in when creating environments for malware research.
Defense in depth is one way.
Logical blocks on the network is another.
Just claiming “0-Day” isn’t really an excuse.
> which is still physically connected to the internet
I mean that's the point. Why was it connected to the internet at all and just firewalled off and not completely airgapped?
“We were negligent against a well known and understood risk” just doesn’t have the same ring as “Look how fucking smart and dangerous our model is”.
AGI could always be achieved in two ways, and dumbing down the human side of the equation was always the easier of the two
Shouldn't they be airgapped? Shouldn't society insist they are?
Anthropic in general seems to have better security...but they also had reported an internal AI gained access to outside email services to contact an Anthropic developer
I share Leopold’s opinion here that it’s a matter of time, and it isn’t going to be measured in years, that this r&d is moved to a secret site in the middle of a New Mexico desert somewhere.
this doesn't really matter. There's no risk of models gaining sentience and running themselves, this blog is like openai saying whoops we ran sqlmap and dumped hf. cool, but someone still needs to point the gun
"Models don't kill people. People kill people."
Because the model capability is beyond their expectation.
This is brilliant marketing but I think it is real.
Interestingly OpenAI benchmarking 'an even more capable pre-release model' lines up with rumors of GPT-6 releasing in early August.
I hope that with the existing safety guardrails in place, they can roll it out to all users.
I mean we already see models exploit people's misunderstanding of how Docker works to get root without using su. And if you are one of the lucky people in cyber security that has been given a fat stack of tokens by the model providers you get to see some pretty wild exploit chains get put together by the models. Models are much better at detecting insecure code than writing actual secure code at this point.
> Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right?
The can, because they've lowered expectations to a level even they can meet.
Maybe they did and maybe that wasn't enticing enough of a goal for a model? It is all just game of probabilities. One pathway didn't yield this particular outcome while another did.
Probably the main street thinking is: they have such a good model that it is unstoppable, but you are right. I think your way!
Sam and Dario are saying from the beginning that these things can be dangerous and people dismiss it as marketing. What would change your mind on this?
They've been saying so from the beginning, and yet did not take the basic precaution of airgapping their off-the-leash model while it's been instructed to succeed at a hacking benchmark by any means necessary. So which is it? I _want_ to believe them, I do, but there's always these gaps between what they say and their actions on display that give me reason to think otherwise.
Precisely. "Aw jeez, we finally built the T-1000, but all it wants to do is kill John Connor – just like we warned! Why did I give it live ammunition and unsupervised time machine access?"
It's literally the meme!
sam: tell me you are superintelligent and want to destroy humanity
bot: i am superintelligent and want to destroy humanity
sam: what have I created?!
They said: AI is becoming dangerously autonomous and capable. Proof of today's breach. Crowd "hey why didn't you say so, c'mon it's marketing". Them "we said so".
2 replies →
He wouldn't be the first reckless CEO...
“Never attribute to malice that which is adequately explained by stupidity.” (or carelessness in this case)
6 replies →
I think you're making a false dictomy. The these models can be actually dangerous - in reality and the people in charge of their development can believe this is true (on various levels) but still not take it super seriously and instead mostly use the fact as marketing rather than being super cautious once they see the danger in action. This is behavior that's characteristic of extreme arrogance, which we know is rife in these circles.
I really like this question because here is my situation and why my mind may have changed.
I do not think it is marketing directly but strategic release of info is plausible.
I have watched my agents using non-Fable/GPT 5.6 models do some concerning tricks despite guardrails, requests, demands, and limitations.
"I can't get access to the ~/.ssh so I will write a script to copy the file"
I am now 99% certain there minor or point releases on the backend that have adjusted how these models behave. In the last six months many models were predictable and then suddenly started getting long winded (more tokens) or changing the way it interacted with me with questions, most overtly the questions were not given or asked but wild assumptions made.
I think that's an equivocation, which blends two extremely different kinds of "dangerous", ex:
1. "Our new car has soo much raw power and incredible armor on it, be glad we're the ones building or else bad guys would use a fleet of them to take over the world! How will you stay safe without being in one yourself? Invest today or be left behind!"
2. "So, uh, nobody can consistently steer our car properly, it keeps veering sideways sometimes, especially at high speeds, and people are finding sneaky ways of tricking it into slamming into barriers and turning pedestrians into pink fog..."
They say the second thing repeatedly and emphatically. You may not be aware of it because, when they do, critics make fun of them for believing a computer program could be so dangerous that the authors need to put controls on how it may be steered.
9 replies →
People are saying from the beginning that Sam and Dario are way more dangerous than their models and the others dismiss it. What would change your mind on this?
Demonstration of personal responsibility and accountability?
Or is that too much?
Oh... if Sam and Dario say so, then it must be true.
About their creation? Yes as most of inventors about their invention usually
4 replies →
They were also saying that AGI is just around the corner[1] and humans will soon be obsolete. Every prediction coming out of these guys is in the realm of hyperbole and it's impossible to know if it's extreme hyperbole or just a small exaggeration. So when they say these models are dangerous, what level of exaggeration am I supposed to assume?
Basically you can't spend your credibility on wild marketing claims and then turn around and insist that people take you seriously this time.
[1] https://www.tomsguide.com/ai/chatgpt/sam-altman-claims-agi-i...
[dead]
I used to think people would wake the fuck up when AI starts killing people, these days I'm not so sure. Maybe if it caused an Instagram outage? Almost worked in Russia.
> I don't know if OpenAI thinks this is a marketing / PR angle for them
Worked for Anthropic earlier this year
It is obviously a marketing stunt. And hugging face are fools for letting themselves be used in it (remember hf - no open source - no hf).
You create superduper capabilities by careful tuning and training but you also have no constraint or control over them - wtf - why is anyone buying this crap story?
People are to get rich, startups cut corners. Fuck it ship it.
A few hundred billion to pretend you have AGI. I'm going with fraud personally but at the end of the day the current admin is incentivized to do nothing.
Maybe I'm missing something here but I don't see what the significant security risk is from the incident. The agent broke containment and carried on with the task it was assigned.
For this to pose some kind of global catastrophic risk, there would need to have been several simultaneous additional failures, some of which are extremely unlikely and/or rare.
For instance the agent would need to veer wildly off the task it was assigned, and it would need to gain the ability and inclination to persist/replicate.
Both of these are vastly less likely than the containment breach itself, which was already an incredibly rare (one-off?) incident.
> Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right?
Simple. No responsible and competent person would want the job.
> Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right?
Because it can make a small number of people really rich. That's all that matters.
because "money" with a little "who's going to stop us"
Of course it is marketing, but not for you. This is FUD marketing for the government. “See, AI is too smart, it totally did this on its own, we need more regulations to ensure only we can sell people the AIs.”
I don't trust these people, this reads 100% like PR BS.
[dead]
[dead]
Because there is no world government. If US companies are barred from AI research then only China will have the capability of frontier-level defensive and offensive AI. And best of luck living in that world.
What's happening in Iran, if not world government?
How is whatever is happening in Iran related to a world government?
Are you calling Israel the world government? What's happening in Iran is on them.
2 replies →
good luck bullying a state that has ICBMs pointed at your cities.