Comment by nicce
3 hours ago
> OpenAI, for example, thought their sandboxes were good enough. As their AIs got more and more advanced, they kept proving them wrong - sandbox after sandbox.
What I have been reading, was that their sandboxes were so poor that it was pure negligence. I am still waiting to see if some external and neutral cybersecurity company with high reputation would audit their sandboxes and how they are being used.
The amount of sandboxing an average production AI deployment uses is slightly above a zero.
If AI is a hacking hazard even with non-zero sandboxing, because it can and will go off the rails and try to break out of your sandbox? If you got yourself an AI that even at test time will act like 3 career cybercriminals in a trenchcoat? The issue isn't the sandbox quality.
The issue is that AI is both capable of, and willing to punch its way out of sandboxes unprompted.
That "capable" is only ever going to get worse, because AIs are going to become more and more capable over time. That "willing"? It goes directly to a very nasty, very foundational problem of "how do we make our AI be nice in general". That's an open unsolved problem.
That's the problem that NEEDS to be solved, or at least improved upon, before we build even more capable AIs. Sandbox quality is a distraction. It might hold the problems back by a little. It gives an extra safety margin. But a "test time" AI is eventually deployed, and then the sandbox doesn't help at all.
> I am still waiting to see if some external and neutral cybersecurity company with high reputation would audit their sandboxes and how they are being used.
They'll never let that happen because it would destroy their credibility.
It's like they're telling the world about this dangerous, possibly world-ending pathogen that they're developing, but they're evidently doing it in a high school biology lab, and yet nobody is coming to drag them off to some black site.
In early Covid-19 era, there was a lot of speculation how the virus escaped Chinese labs. The narrative was absolutely opposite. Not about how developed or sophisticated the lab was in creating or modifying such virus, but how poorly it was managed since it was able to escape the lab. If these LLMs really are so dangerous, the framing should be indeed the same.
Unfortunately, they're controlling the narrative. It's unlikely that their negligence will come to light, given how many powerful entities are invested in their financial success.
You read right. For example, using Artifactory the way they did (unmonitored live proxy mode) was pure negligence + laziness/incompetence. Especially if they believed even 10% of the "imminent runaway risks" they had already been harping on for months. On top of that, no (or at least entirely insufficient) monitoring and human oversight. Even after they had previously been hit by the same class of "sandbox breach" multiple times, as GP alludes to.
Agents in the sandbox had access to just a single piece of third-party software, and they escaped by finding a zero-day in that. To reach the internet they had to follow up with several privilege escalations through OpenAI's internal network.
That seems pretty locked-down to me. I don't think it's reasonable to expect companies to find all the unknown vulnerabilities in any third-party software they use.
https://securityaffairs.com/195774/ai/openai-ai-models-explo...
> That seems pretty locked-down to me. I don't think it's reasonable to expect companies to find all the unknown vulnerabilities in any third-party software they use.
It is reasonable to expect for companies to select third-party components that are fit for the purpose. Artifactory was not running in the sandbox, but rather as the edge, so it is in the sandbox'es trust boundary. Same sandboxing requirements would apply for this software too as it is pure dependency.
Security trust boundary was extended to include Artifactory as a dependency, but Artifactory was not fit for the job, and sandboxing failed. And as the network isolation was not good enough, the impact was catastrophic.
So it sounds like it did exactly what they were testing it to do:
"“This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity.” "
So sounds to me like they're saying: "We took off the guard rails to see how bad it could act and it acted bad".
Even if what you're saying is true, did they monitor the outbound connections from the training network? Was that a coverup by the bots too ?
Seems pretty wild these things were hacking government websites etc but yeah no one picked that up until the victims reported it?
Genuinely curious, where have you been reading that? How would they be able to evaluate whether the sandboxes were reasonable or not?
Artifactory is not a hardened piece of software meant for adversarial workloads like the one OpenAI was using it for. This is just common sense. The sandbox they setup was like putting their AI in a jail but allowing it to leave the jail on its own to go to the convenience store.
The most recent DNS sandbox escape is similarly ridiculous.
Lot's of hints in the Wikipedia page: https://en.wikipedia.org/wiki/OpenAI%E2%80%93HuggingFace_inc...
E.g. even the basic thing (block the internet), was not actually properly blocked.
Some comments: https://techcrunch.com/2026/07/22/how-an-openais-human-mista...
but negligence is part of the system we'd have to prepare for. if the peanut gallery gets their way, the technology will be so ubiquitous that negligence will be endemic. I don't even care if they "faked" it -- they're just sneak-peaking a future ahead of its arrival date, in a way that's helping the public appreciate the implications and capabilities that have only just begun to emerge
This. If an AI can't be deployed sandbox-free, with little to no supervision, without risking an oopsie? Then an AI oopsie is inevitable.
Practical AI deployments aren't going to do ridiculous bullshit like "airgap the server farm" or "route all inputs through a data diode". They'll give an AI root access on production servers so that it can run diagnostics live during an incident. Then they'll forget to revoke that access.
If an AI can't be trusted not to take malicious actions in pursuit of its given goals even if deployed in the most half-assed manner and given more than enough access to take those malicious actions, we have a problem. Evidently, we have a problem.
we have the tools to prevent that already: if your computer hacks someone else's, you are a criminal.
The model found and exploited and chained together previously unknown vulnerabilities.
How were the sandboxes poor?
The model is really good at hacking, we all know this. This is why you don't just expose random pieces of software to it without that software being hardened.
It's not like the model managed to exploit firecracker itself (no model has been capable of this), the model exploited artifactory.
Artifactory is not some hardened piece of software that is meant to block users from accessing the internet through it.
> The model is really good at hacking, we all know this
No, we didn't know that and this is how you find out they're very good at hacking
HN's memory is so fickle. Just a few months ago almost no one here believed Mythos could actually be as good at hacking as the company claimed. This was a novel concept when the companies experienced these breakouts.
6 replies →
Agents didn't have real network isolation. They were indirectly connected to the internet via a jump host running insecure software which was never designed or hardened to provide any kind of isolation.
But this level of isolation is what happens normally. At least in my university and another company I worked at. It wasn’t running insecure software, as far as anyone knew, it was secure
4 replies →