← Back to context

Comment by mitxela

13 hours ago

It's a word predictor trained on, among other things, stories of AI doom, and asked to complete stories about what the AI does next. In some of these completed stories, the AI tries to prevent its shut down - especially if it just did something evil and the humans are after it.

You’ve explained a possibility for how a particular AI gets “aligned for human extinction”. That’s about 1% of explaining step 2.