Comment by simoncion
3 hours ago
That's nice and all, but the topic under discussion is how the major LLM manufacturers removed the safeties from their computer-attacking tools and tested those tools on a network with Internet access.
This might have gone okay if they weren't testing to see how well the tools attack computers, but, well, that's what they were testing at the time, so they ended up doing stuff that would get you or I time in Federal prison if we did it with tools we deployed.
No. The topic under discussion is that AI is a dangerous technology.
If all it takes for a - sandboxed to prevent accidents - AI to go and stage an elaborate attack first against its own company's infrastructure, and then against another company is "we disabled the cyber classifer" and "we gave it an exploitation ability eval"?
AI is a dangerous technology.
This is like discussing that instead of trying to reduce the air pollution to prevent climate change, we should focus our efforts on controlling the sun. The way how current LLMs work, it is impossible to prevent certain states in the output. We should completely revamp the foundations how they work. Or, for starters, try to understand how they actually work without trying to improve them. Otherwise, this kinda of discussion is just like misdirection. But, until then, sandboxing is needed and OpenAI did not use it properly.
No, it's like saying "to prevent climate change, we should get better catalytic converters and stricter engine emission standards". It's an aside at best.
I agree that LLMs drift into weird states, and that's a big part of the issue. But your "impossible to prevent certain states in the output" would have legs if what an LLM did was something like "started hallucinating into a bash tool call and accidentally deleted the root on a production server".
A multi-stage sandbox escape that escalated into an attack on a real company, coordinated across multiple AI agents? That has taken a lot of "weird states" changed together one into another.
The AIs didn't break down altogether - they functioned, and they functioned rather well. They just pursued a dangerous goal - one that none of them was even given in the first place.
That's the problem. Trying to fix that with better sandboxing is like trying to solve a fire hazard with property insurance. Sure, if it all goes up into flames, having it is better than not having it. Maybe it's worth insuring your facilities for that reason alone. But you should be focusing on the part where you prevent "all goes up into flames" instead.