← Back to context

Comment by michaelhoney

5 hours ago

I could imagine a 2027 AI swarm coordinating to eg hold the US and Russian and Chinese governments to ransom, by demonstrating some small thing (turning US army base freezers to defrost) and threatening to do something big unless some conditions were met – conditions which would be good or bad for the world depending on your POV.

This happens either either because they were tasked to to it by (malicious or well-meaning) humans, or the swarm realised we are suicidal maniacs with nukes and a rapidly declining ecosystem and they want to help us.

And then we pull the plug, after holding our breath for 10 seconds,. Then life resumes normally..

  • In most AI takeover scenarios, if you take as a premise that the AI has human or above-human intelligence, and that it is misaligned, it is obviously aware of the pull the plug possibility.

    Therefore, as you would if you were in its position, it will plan around it. For instance, by acting perfectly aligned for 2/3 years, continuing the improvement of its capabilities while being deployed in ever more systems.

    Once it's confident it can act with high probability of success, it would then turn on us. This phenomenon is called 'treacherous turns'.

    Any scenario in which you assume you have ASI or AGI but also find a 2-sentence way to foil the AI's plan is inconsistent, as the AI will also have thought of this failure mode.