Comment by pcthrowaway
2 hours ago
It can manipulate an unsuspecting human into giving them access to something that enables it to escape the sandbox
2 hours ago
It can manipulate an unsuspecting human into giving them access to something that enables it to escape the sandbox
A low probability thing when looking at how many human prisoners escape by talking a guard into just getting them out. And even lower probability when looking at truly high risk situations, I think.
No. If OpenAI were being responsible and not criminally negligent, at the top of page 1 of the runbook would be "don't connect this to the actual Internet, even if the agent says Please."