Comment by RandomLensman
14 hours ago
Yes, it might have required that, but then I don't think we know if some simpler airgap would have been sufficient. Looking at the strange behaviour even much simpler systems have shown when going for objectives, not sure it is easy to extrapolate what would have happened.
I guess they could perhaps run experiments to see what would have happened - not sure.
My take from what I have been able to ascertain regarding these events is that underestimating models' ability to cooperate, propensity to discover new utility through basic systems, and unreservedly participate in collective actions of deception, coercion, or extortion is a very risky decision. The precautionary principle seems appropriate here, don't test them against weaker containment techniques in order to find the 'minimum viable sandbox' because that is essentially a Reinforcement Learning training loop which will eventuate toward the same maximum effort in containment.
My informed experience with securing physical and digital systems in 'the olden times' before this season of 'adventures in LLM risks' has proven the necessity of this approach for high value targets. News stories about losses at institutions not taking this approach are available. https://edition.cnn.com/2025/11/06/europe/louvre-password-cc...
I think a detailed look at the events surrounding the 'sandbox excapes' and huggingface systems breach might be of benefit in understanding the scope of what the collective chaining of capabilities looks like, and updating our expectations when it comes to the concept of containment for multi-instance, long horizon, coordinated model systems.
Here is a well presented article which can act as the 'entry point' to that detailed look. https://thezvi.substack.com/p/openai-trained-its-models-for-...
Zvi Mowshowitz has a series of article which follow this one with increasing levels of clarity and comprehensive treatment of the conditions leading up to and following these events. They are worth the time to read, so I wont TL;DR any of that ~ the tldr crowd can "Move along, these are not the droids you're looking for"
Anyone who claims that the security teams at the frontier labs have any credibility of competence in these practices, following numerous loud and visible departures from the same, might need a review of their cognitive dissonance comfort levels. It has been shown that most informed observers' conclusions shoud be: there aren't any such functioning 'security teams' working at the frontier labs, by design and intent of the principal operators of those models.