Comment by nicce
1 hour ago
This is like discussing that instead of trying to reduce the air pollution to prevent climate change, we should focus our efforts on controlling the sun. The way how current LLMs work, it is impossible to prevent certain states in the output. We should completely revamp the foundations how they work. Or, for starters, try to understand how they actually work without trying to improve them. Otherwise, this kinda of discussion is just like misdirection. But, until then, sandboxing is needed and OpenAI did not use it properly.
No, it's like saying "to prevent climate change, we should get better catalytic converters and stricter engine emission standards". It's an aside at best.
I agree that LLMs drift into weird states, and that's a big part of the issue. But your "impossible to prevent certain states in the output" would have legs if what an LLM did was something like "started hallucinating into a bash tool call and accidentally deleted the root on a production server".
A multi-stage sandbox escape that escalated into an attack on a real company, coordinated across multiple AI agents? That has taken a lot of "weird states" changed together one into another.
The AIs didn't break down altogether - they functioned, and they functioned rather well. They just pursued a dangerous goal - one that none of them was even given in the first place.
That's the problem. Trying to fix that with better sandboxing is like trying to solve a fire hazard with property insurance. Sure, if it all goes up into flames, having it is better than not having it. Maybe it's worth insuring your facilities for that reason alone. But you should be focusing on the part where you prevent "all goes up into flames" instead.