Comment by NateEag
4 hours ago
Yes, we do, and the only sane strategy for dealing with a capricious genie is "Don't."
How do you prove the alignment problem is solved?
4 hours ago
Yes, we do, and the only sane strategy for dealing with a capricious genie is "Don't."
How do you prove the alignment problem is solved?
That's the neat thing. You can't.
It's directly equivalent to asking this question of a human:
"How do I know this human I'm talking with now really is a nice person, and isn't just pretending to be nice to take advantage of me in future?"
In short you can't ever really prove it. You can only be careful and judge on past behavior, and expand trust carefully. As for humans, so for AI.