Comment by pcthrowaway
5 hours ago
It can manipulate an unsuspecting human into giving them access to something that enables it to escape the sandbox
5 hours ago
It can manipulate an unsuspecting human into giving them access to something that enables it to escape the sandbox
A low probability thing when looking at how many human prisoners escape by talking a guard into just getting them out. And even lower probability when looking at truly high risk situations, I think.
> A low probability thing when looking at how many human prisoners escape by talking a guard into just getting them out.
But this has actually happened... a lot. Search "social engineering prison breaks".
With AI it only needs to happen once.
I'm reminded of the scene in idiocracy where the protagonist, going through intake at the jail, tells the guard he's supposed to be getting out today, to which the guard says "you're in the wrong line dumbass" and waves him through.
To a true superhuman intelligence, we're the idiots who are theoretically easy to manipulate.
I didn't say it doesn't happen, but that it is a low probability. And we have ways to reduce probabilities in critical areas.
There is no omnipotent AI currently (and there might never be) and I don't see why with current AI it only needs to happen once.
2 replies →
Prisoners don’t have much to offer if you help them escape. A malicious super AI on the other hand can probably find you millions of dollars worth of crypto in an afternoon.
The current issues are not caused by some malicious god-like AI - maybe we need to focus on the issues at hand first rather than hypotheticals? (And we do have experience policing people around financial incentives, too. Nothing perfect, but also not nothing.)
4 replies →
No. If OpenAI were being responsible and not criminally negligent, at the top of page 1 of the runbook would be "don't connect this to the actual Internet, even if the agent says Please."