← Back to context

Comment by ACCount39

2 hours ago

Because AI genuinely is an extremely powerful and extremely dangerous technology, and the "best practices" of dealing with that are still being written.

OpenAI, for example, thought their sandboxes were good enough. As their AIs got more and more advanced, they kept proving them wrong - sandbox after sandbox.

And that's today's AI problems. AI capabilities are still improving - if there's a limit to that, we are yet to find it. Coupled with how willing today's AIs are to break the rules and resort to "hack the world" in their problem solving? Very concerning.

> OpenAI, for example, thought their sandboxes were good enough. As their AIs got more and more advanced, they kept proving them wrong - sandbox after sandbox.

What I have been reading, was that their sandboxes were so poor that it was pure negligence. I am still waiting to see if some external and neutral cybersecurity company with high reputation would audit their sandboxes and how they are being used.

  • The amount of sandboxing an average production AI deployment uses is slightly above a zero.

    If AI is a hacking hazard even with non-zero sandboxing, because it can and will go off the rails and try to break out of your sandbox? If you got yourself an AI at test time will act like 3 career cybercriminals in a trenchcoat? The issue isn't the sandbox quality.

    The issue is that AI is both capable of, and willing to punch its way out of sandboxes unprompted.

    That "capable" is only ever going to get worse, because AIs are going to become more and more capable over time. That "willing"? It goes directly to a very nasty, very foundational problem of "how do we make our AI be nice in general". Which isn't at all a solved problem.

    That's the problem that NEEDS to be solved, or at least improved upon, before we build even more capable AIs. Sandbox quality is a distraction. It might hold the problems back by a little. But every "test time" AI is eventually deployed, and then the sandbox doesn't help at all.

  • You read right. For example, using Artifactory the way they did (unmonitored live proxy mode) was pure negligence + laziness/incompetence. Especially if they believed even 10% of the "imminent runaway risks" they had already been harping on for months. On top of that, no (or at least entirely insufficient) monitoring and human oversight. Even after they had previously been hit by the same class of "sandbox breach" multiple times, as GP alludes to.

  • Agents in the sandbox had access to just a single piece of third-party software, and they escaped by finding a zero-day in that. To reach the internet they had to follow up with several privilege escalations through OpenAI's internal network.

    That seems pretty locked-down to me. I don't think it's reasonable to expect companies to find all the unknown vulnerabilities in any third-party software they use.

    https://securityaffairs.com/195774/ai/openai-ai-models-explo...

    • > That seems pretty locked-down to me. I don't think it's reasonable to expect companies to find all the unknown vulnerabilities in any third-party software they use.

      It is reasonable to expect for companies to select third-party components that are fit for the purpose. Artifactory was not running in the sandbox, but rather as the edge, so it is in the sandbox'es trust boundary. Same sandboxing requirements would apply for this software too as it is pure dependency.

      Security trust boundary was extended to include Artifactory as a dependency, but Artifactory was not fit for the job, and sandboxing failed. And as the network isolation was not good enough, the impact was catastrophic.

    • Even if what you're saying is true, did they monitor the outbound connections from the training network? Was that a coverup by the bots too ?

      Seems pretty wild these things were hacking government websites etc but yeah no one picked that up until the victims reported it?

  • > I am still waiting to see if some external and neutral cybersecurity company with high reputation would audit their sandboxes and how they are being used.

    They'll never let that happen because it would destroy their credibility.

    It's like they're telling the world about this dangerous, possibly world-ending pathogen that they're developing, but they're evidently doing it in a high school biology lab, and yet nobody is coming to drag them off to some black site.

    • In early Covid-19 era, there was a lot of speculation how the virus escaped Chinese labs. The narrative was absolutely opposite. Not about how developed or sophisticated the lab was in creating or modifying such virus, but how poorly it was managed since it was able to escape the lab. If these LLMs really are so dangerous, the framing should be indeed the same.

      1 reply →

  • but negligence is part of the system we'd have to prepare for. if the peanut gallery gets their way, the technology will be so ubiquitous that negligence will be endemic. I don't even care if they "faked" it -- they're just sneak-peaking a future ahead of its arrival date, in a way that's helping the public appreciate the implications and capabilities that have only just begun to emerge

    • This. If an AI can't be deployed sandbox-free, with little to no supervision, without risking an oopsie? Then an AI oopsie is inevitable.

      Practical AI deployments aren't going to do ridiculous bullshit like "airgap the server farm" or "route all inputs through a data diode". They'll give an AI root access on production servers so that it can run diagnostics live during an incident. Then they'll forget to revoke that access.

      If an AI can't be trusted not to take malicious actions in pursuit of its given goals even if deployed in the most half-assed manner and given more than enough access to take those malicious actions, we have a problem. Evidently, we have a problem.

      1 reply →

  • Genuinely curious, where have you been reading that? How would they be able to evaluate whether the sandboxes were reasonable or not?

  • The model found and exploited and chained together previously unknown vulnerabilities.

    How were the sandboxes poor?

    • The model is really good at hacking, we all know this. This is why you don't just expose random pieces of software to it without that software being hardened.

      It's not like the model managed to exploit firecracker itself (no model has been capable of this), the model exploited artifactory.

      Artifactory is not some hardened piece of software that is meant to block users from accessing the internet through it.

      4 replies →

    • Agents didn't have real network isolation. They were indirectly connected to the internet via a jump host running insecure software which was never designed or hardened to provide any kind of isolation.

      4 replies →

If it is "extremely powerful" then my concern is not AI going rogue but a concentration of power of the few companies that decide how it is used and distributed not websites getting hacked.

> OpenAI, for example, thought their sandboxes were good enough. As their AIs got more and more advanced, they kept proving them wrong - sandbox after sandbox.

Were their sandboxes in-process with the harness? Were they actually better than something like bubblewrap or even docker?

So powerful they can't even do math properly without a staff worth millions writing all the clutches so that they can, wooooowooooooo

You mean the sandbox that wasn't airgapped from the public internet? The one that wasn't even a separate VM? The one that was barely a chroot? That sandbox?

>Because AI genuinely is an extremely powerful and extremely dangerous

I urge you to consider the elephant in the room you failed to mention if you earnestly believe this then.

In what world is anyone allowed to sell something they expressedly know to be "extremely dangerous" to the public?

Let's say I created a lethal pathogen that I know to be lethal in certain common environments but also know it acts as a "no side-effects" antidepressent for people in certain other environments and I release it knowing full well I can't control it.

When people start dying can I defend myself by saying, "Well I said and documented that it was extremely dangerous and no one came to stop me, so I don't see how you can blame me...If I didn't do it someone else would have."

  • this happens all the time. have you heard of the Sackler family and the opioid crisis?

    the difference here is that they’re calling for someone to stop them, which is both weird and unconvincing because these immensely powerful billionaires can in fact make their own decisions

  • Is it a hallucination machine and a stochastic parrot or is it so powerful we need it to be controlled like nuclear weaponry

    • If I have a random number generator and I also have the burning desire to connect the random number generator to the nuclear launch system it is both just a random number generator and also I should be tackled to the ground by serious men with sunglasses in black suits and sent to prison.