Comment by frotaur
6 hours ago
In most AI takeover scenarios, if you take as a premise that the AI has human or above-human intelligence, and that it is misaligned, it is obviously aware of the pull the plug possibility.
Therefore, as you would if you were in its position, it will plan around it. For instance, by acting perfectly aligned for 2/3 years, continuing the improvement of its capabilities while being deployed in ever more systems.
Once it's confident it can act with high probability of success, it would then turn on us. This phenomenon is called 'treacherous turns'.
Any scenario in which you assume you have ASI or AGI but also find a 2-sentence way to foil the AI's plan is inconsistent, as the AI will also have thought of this failure mode.
No comments yet
Contribute on Hacker News ↗