Comment by SaucyWrong
8 hours ago
> while replicating wildly
Earnest question: by what mechanism that exists today would the achieve that in a way humans on top top of the situation could not curtail?
All of this runs on top of compute in meatspace that humans can disconnect.
I am not an AI super mind hell bent on consolidating my power by leveraging chaos to take control of humanity’s resources, but if I were then sending one million deepfaked ransom emails to impressionable people would be the best tool for effecting change in meatspace.
We have your daughter / dog / Amazon delivery. If you ever want to see her / him / it again, plug this USB drive into the control panel at your station / let off the parking brake roll your car into this substation / change the meatpacking thermometers to read 8C lower than calibrated / ground your vessel on this sandbank / send an envelope of white powder to these addresses / set fire to the following hospitals / …
I'm waiting for some data center to be built where no one can agree on who actually commissioned and payed for the thing. Every body "just followed orders" until it turns out that it was Grok.
Imagine you're the AI. Give yourself a solid minute to brainstorm ideas.
Here's my answer, as a non-superintelligent human: "see to it that the humans on top of the situation have a compelling financial interest in the systems not disconnecting".
In nuclear engineering, where safety is taken seriously, it's not enough to end the conversation at "the humans in charge can always simply shut down the reactor during a meltdown" or "a meltdown has never happened before, so we don't have to design safety systems before one does".
The reason nuclear reactors are dangerous is because if you turn off the power cooling them down, they react (and radiate) more.
If you turn off the power cooling a data center, the servers within rapidly stop doing any computing.
Positive feedback loops are dangerous. Negative ones self-regulate.
Yes. But nobody is worried about datacenters overheating and physically exploding, so I'm not sure what comfort that's supposed to provide? The positive feedback loops in AI operate at different levels than that, but they deserve safety engineering all the same.
For example, if the head of cyber security at your company suggested there's no need to worry about hacker infiltration or worms because one can always unplug one's computer as the primary defense mechanism, you might find that a little lacking. Will you be able to unplug the computer before the damage is done? Will it spread to other systems before you detect it? How will you unplug the computer if the attack is from an external facility? What if an attack happens but the boss says the computers have to keep running because an important customer is monitoring uptime? What if the attack goes unnoticed because it looks like a benign service?
Now imagine the head of cyber security answers by saying "actually you don't even need to unplug them, you can just wait for the computers to overheat, thus solving all concerns."
I can imagine small snippets of malware-like code that behave like a virus, using a host’s LLM/AI to self-edit/evolve its payload.