Comment by includenotfound

8 hours ago

The "sandbox" they used was apparently made of thin paper exposed under a day of heavy rain, too. You'd think, if they truly believed the model is so dangerous, they'd run it in a VM without a network adapter.

I brought this up to someone else and was told that airgapping is apparently much more expensive than I'd naively think.

I still think this is a sign that they are not taking their own rhetoric seriously.

  • it's really weird to hear frontier labs say "our internal models are basically AGI" while also saying "airgapping is too hard uwu".

    if your internal models are so damn good, they should be able to "one shot" airgapping... right?

  • Agents need packages like the rest of us. Ruby gems, npm packages, Maven, pip, docker images..

    Not surprised this is always what they have and hack.

    Who would use an Agent that spends $10,000 re-implementing some OAuth lib or reverse-engineering a proprietary lib when it's free on the internet?

    • You don't need a full air gap. Set up a microVM with network access limited to local network and send all package requests through a filtering gateway that only allows normal download endpoints. Or self host a big collection of popular packages if you need extra security.

      3 replies →

    • There was and continues to be no reason to share the package manager between models. This was begging for abuse.

    • > Agents need packages like the rest of us. Ruby gems, npm packages, Maven, pip, docker images..

      Yes, yes they do, but read through artifact proxies are dodgy as fuck, which is why and facebook (and I assume a fuckload others) don't have them.

      Also semi-airgapped labs are a lot less expensive than you think at that scale. Once you have to do multi-region VLANs with machine certs before you get access to juicy VLANs, the difference between "no internet for you" and "mostly airgapped" falls to almost zero.

      Also I would want an artifact mirror because a) that give a good signal about how the model reacts, and what training material its latched onto, b) it hides what the models are doing from the outside.

  • It's expensive if it wasn't part of the planning and design. The same as 'security' is expensive, or compliance with regulations is expensive.

    It is also a choice to not do any or all of the above.

> You'd think, if they truly believed the model is so dangerous...

They would have been watching what it does, especially when running it on ExploitGym of all benchmarks... that is criminal worthy neglegence

yes, that is the kicker

These same people who supposedly believe these agents pose an existential threat to humanity apparently fired up 10,000 of them and left them unsupervised for weeks.