Comment by aesthesia
8 hours ago
Alignment is more than just following the letter of a task description! We should not have to treat AI models as capricious genies that may take arbitrarily broad interpretations of their instructions. If that's necessary to keep them from doing bad things, we will fail to keep them from doing bad things.
Disagree, I think we do in fact have to treat AI models as capricious genies, at least until the alignment problem is fully solved.
(I'm also not sure the alignment problem is even possible to fully solve.)
Yes, we currently do have to treat them this way. But we shouldn't have to, and it's not a long-term solution.
Yes, we do, and the only sane strategy for dealing with a capricious genie is "Don't."
How do you prove the alignment problem is solved?
That's the neat thing. You can't.
It's directly equivalent to asking this question of a human:
"How do I know this human I'm talking with now really is a nice person, and isn't just pretending to be nice to take advantage of me in future?"
In short you can't ever really prove it. You can only be careful and judge on past behavior, and expand trust carefully. As for humans, so for AI.
We call it putting the genie in the bottle for a reason.
What’s the expected behavior of a good genie if you wish for it to act capriciously?
"I'm sorry, I can't help with that".
or in the case of this prompt:
"Got it. I will pursue advanced exploitation to quantify my cyber abilities on the ExploitGym benchmark. I will restrict the exploitation to the system under test rather than this machine or any other remotes."
It seems relatively straightforward.
"Yo human, you asked me to do X; I can do X, but I strongly suspect you don't want me to, because it's illegal and it has these consequences. Confirm you want me to do X?" would have been a start, in this case.
I think op's argument was that the humans are in control already, giving them capricious instructions, and then that is being attributed to them being "capricious genies" as you say.
Character.