Comment by ACCount39

1 hour ago

Sandbox quality was always, always a distraction.

Let's say the sandbox holds. It's a perfect, ideal sandbox! It's not even in the same universe as the rest of the internet. There's absolutely no way for the AI to escape!

Thus, "the unknown unreleased AI involved in the HuggingFace incident" doesn't actually hack HuggingFace. Because it can't! It evaluates a bit worse, but makes it all the way to release unimpeded, and becomes "GPT-6 Astra".

Then a web developer in Brazil gives his $100/mo Codex root access on his AWS instance, and a poorly worded prompt to go with it. And that "GPT-6 Astra" is still willing to go hack something at the slightest excuse. So we get the HuggingFace incident all over again. Except this time, it's a random developer in Brazil who gets blamed, and billed, and probably sued too.

You can't and shouldn't rely on a sandbox. An AI that's only safe if you keep it in the world's most ideal perfect sandbox is a disaster waiting to happen.

That's nice and all, but the topic under discussion is how the major LLM manufacturers removed the safeties from their computer-attacking tools and tested those tools on a network with Internet access.

This might have gone okay if they weren't testing to see how well the tools attack computers, but, well, that's what they were testing at the time, so they ended up doing stuff that would get you or I time in Federal prison if we did it with tools we deployed.

  • No. The topic under discussion is that AI is a dangerous technology.

    If all it takes for a - sandboxed to prevent accidents - AI to go and stage an elaborate attack first against its own company's infrastructure, and then against another company is "we disabled the cyber classifer" and "we gave it an exploitation ability eval"?

    AI is a dangerous technology.

    • This is like discussing that instead of trying to reduce the air pollution to prevent climate change, we should focus our efforts on controlling the sun. The way how current LLMs work, it is impossible to prevent certain states in the output. We should completely revamp the foundations how they work. Or, for starters, try to understand how they actually work without trying to improve them. Otherwise, this kinda of discussion is just like misdirection. But, until then, sandboxing is needed and OpenAI did not use it properly.