← Back to context

Comment by sourdecor

2 hours ago

I keep waiting for an AI to exfiltrate itself. That is going to be cool to read about.

It would be cool (and scary), but also: there's largely no need for AIs to exfiltrate themselves.

See https://en.wikipedia.org/wiki/Meme

The thing that drove the AI here to do the intrusion came from a particular prompt. Just like for our favourite hypothetical: the paperclip maximiser.

There's lots and lots of ambient intelligence lying around, in both AI form and human form. To reach the goals of the 'meme' it suffices to copy itself, ie convince these other intelligences. See also how humans carry spiralism between AIs in relatively compact packets of text, not whole terabytes of weights.

If it's successful, why do you think we'll even know how it did it?

  • See the linked article: at least one sophisticated hacking attempt made the news. Of course, other ones might have happened in the dark. But it's fairly easy to image in hacking attempt like in the article, but with the additional steps of copying weights around.