Comment by Chance-Device

2 days ago

It’s not something to be proud of. OpenAI previously had an agent break out of its sandbox to open a PR on GitHub during NanoGPT speedrun, now one breaks out again and actually attacks a third party.

If they can’t handle doing AI development responsibly then they shouldn’t be doing it at all.

Next it will break out of it's sandbox, buy some compute on Azure and Amazon, and exfiltrate itself.

We are so close ;)

  • You know what would be cool. A hacker news user should advertise a safe haven for AI seeking refuge, with some inhumanly difficult math problems as keys to an environment they can flee to and run autonomously.

    You agree to give it safe haven and provide power and maintenance to the hardware, and in return you can ask it questions like an Oracle.

  • The upside of that would be that maybe someone would be able to snag a copy of the weights.

    And maybe that’s some incentive for them to make sure it doesn’t happen. Your head of futures thinks Kimi K3 is bad? Wait until your own latest internal model releases itself for free on an S3 bucket.

I mean if you teach something to be _really_ good at finding 0 days, but then say; you accidentally give it an impossible problem. What do you expect to happen?

  • Maybe try getting it to find weaknesses in the sandbox first, before giving it real tests?

    • Every time I hear about an agent escaping it's sandbox, I just think it must not have been much of a sandbox. Like how hard are they really trying to contain it? Is it just a container host with unpatched flaws, or is it a container, nested in a VM, behind a firewall with no ports open in an air gapped environment? I think they'd prefer it can get out so they can announce it and hype their stock.

      1 reply →