Comment by ACCount39
1 hour ago
No, it's like saying "to prevent climate change, we should get better catalytic converters and stricter engine emission standards". It's an aside at best.
I agree that LLMs drift into weird states, and that's a big part of the issue. But your "impossible to prevent certain states in the output" would have legs if what an LLM did was something like "started hallucinating into a bash tool call and accidentally deleted the root on a production server".
A multi-stage sandbox escape that escalated into an attack on a real company, coordinated across multiple AI agents? That has taken a lot of "weird states" changed together one into another.
The AIs didn't break down altogether - they functioned, and they functioned rather well. They just pursued a dangerous goal - one that none of them was even given in the first place.
That's the problem. Trying to fix that with better sandboxing is like trying to solve a fire hazard with property insurance. Sure, if it all goes up into flames, having it is better than not having it. Maybe it's worth insuring your facilities for that reason alone. But you should be focusing on the part where you prevent "all goes up into flames" instead.
No comments yet
Contribute on Hacker News ↗