← Back to context

Comment by zahlman

21 hours ago

It's been shown that people could be convinced to say "okay I'll let you out of the box". That doesn't mean that the person thus convinced is actually capable of doing so.

Of course, there are huge risks there. But this goes more towards explaining the fact that OpenAI's experiments thus far have worked the way they did, than it does towards actually informing a useful threat model for OpenAI to follow.