← Back to context

Comment by ball_of_lint

1 day ago

Two things can be true. Yes this appears to be negligence on the part of OpenAI.

However, making a secure 'sandbox' is quite hard. There's a huge variety of exploits that exist today, including many we don't know about. Strong models have already shown a capability of finding and using such bugs.

Even one of the strongest boxes we can imagine, literally just a text interface a human can read, has been repeatedly shown to allow unfriendly AI to escape containment: https://www.lesswrong.com/w/ai-boxing-containment

Hadn't seen that link. Thanks!

Protecting against an AI convincing the human to let it out of the box is an interesting challenge. Even a dual-keyed system to do so only requires convincing two people, although I suppose more elaborate containment mechanisms can be devised. N-keys is democracy and that doesn't seem immune.

Two random thoughts:

* This box we want to keep AIs in reminds me a bit of the story of Pandora's box

* I'm also reminded of the original alignment problem in Genesis 3 where one creature convinces another to take an unaligned action.

Containment is a pretty fundamental tricky security problem actually once persuasion is part of the threat model.

This also happens in prisons where inmates are able to manipulate guards or nurses.

e.g.

- guard is chewing gum

- inmate says "where's my piece of gum?"

- this is b/c it's against the rules to chew gum and the inmate is implicitly stating this

- the guard should go to his boss and admit he made a mistake

- but decides to give the prisoner a piece of gum instead

- the inmate now has leverage over the guard b/c the guard broke the rules and then "covered it up" by giving the inmate gum

- this can then get slowly escalated into bigger and bigger asks from the inmate until the guard is bringing in drugs

This is not hypothetical. There are documented cases of both this and male prisoners "seducing" multiple female prison employees despite their being explicit warnings about this.

Both of the above are from this book: Maximum Insecurity: A Doctor in the Supermax by William Wright [0]

I highly recommend it for both the stories above but also a look inside prisons in general and how they operate with medical care more specifically.

0 - https://amzn.to/3TVhGBV

It's been shown that people could be convinced to say "okay I'll let you out of the box". That doesn't mean that the person thus convinced is actually capable of doing so.

Of course, there are huge risks there. But this goes more towards explaining the fact that OpenAI's experiments thus far have worked the way they did, than it does towards actually informing a useful threat model for OpenAI to follow.

I find the super-persuasion argument annoying. Yes, an AI could socially engineer a human to give it what it wants. However, the vast majority of AI conversations in OpenAI's training runs were unmonitored. The agents in the Hugging Face hack correctly understood that human supervision was nonexistent and acted accordingly. I suspect the agents in this breach did the same.

All the AI-boxing arguments from the LessWrong people assume that there will be one AI, with one chat context, in conversation with a human. The actual deployment environment is more like many contexts talking to themselves, so arguments about human supervision are limited. Technical constraints and IT security practice are what matters in the actual deployment environment.

  • It's what we know mattered for this model. It doesn't tell us about future models; we already know this strategy is unsound, so it is silly to rely on it.

>has been repeatedly shown to allow unfriendly AI to escape containment:

Has it been shown, or has the cult simply updated their tenets to preclude AI containment?

I swear any idiot with the tiniest bit of security or networking experience could box up an LLM. Failing to do it properly is a choice.

  • It’s beyond absurdity at this point, but the marketing plan of “describe our poor sandbox design as powerful agent capabilities” clearly works.