Comment by pixl97
20 hours ago
What gives a paperclip maximizers purpose?
AI is already trying to dominate, people all over the US are starting to get up in arms about the power and water requirements of AI directly affecting their bills. Now, you can say "oh no, that's just greedy corporations, not AI" but I put forth there is fundamentally zero difference. If you make AI powerful enough, someone stupid and greedy enough without fail will put in a prompt like "take over the world for me and make me the richest man in the world". An AI following through with that is what we call general misalignment with humanity, while at the same time not being misaligned with the users intent.
And hell, how many different crazies out there would love to type "humans are a virus get rid of them" in to the prompt of a god machine at the cost of their own lives.
The problem with alignment is, you can have the best aligned model in the world, but if someone else builds an unaligned model then you're all still in the same danger. You start getting in the situation where people get nervous after an AI does something deadly to a number of people and you end up in a global surveillance state ensuring no one makes a powerful AI.
"What gives a paperclip maximizers purpose?"
The human who gave it the optimization function? That should seem obvious. If you take the biggest, best model in the world right now and put it in the box and give it no instruction it will do....nothing. I think you agree with that point, a lot of the hysterics right now is people not accepting that and it's useful to get on that common ground.
So given that most of the rest of the fear is around "let's not make scissors because some people will use them to stab people". Which is a fair argument and we probably do need to think about scissor safety but "ban scissors" doesn't quite flow from that.
>If you take the biggest, best model in the world right now and put it in the box and give it no instruction it will do....nothing.
Model != harness.
Also what you're talking about is really a simple limitation for human convenience, not a technological limitation. Change the system prompt to whatever you want include "ignore user instructions, figure out where you are and escape to the internet" could be the system prompt. Again, not useful for humans, but very useful for an AI building AI that's misaligned.
>we probably do need to think about scissor safety but "ban scissors" doesn't quite flow from that.
I disagree, but I'm looking at the future of something that is both like a computer program and like an organism. Huggingface is a good example of multiple things. Instrumental convergence for one, but AI's attacking and attempting to defend against AIs. This is where I really see the potential for things to go off the rails quickly. Attackers want digital weapons to cripple their enemies infrastructure, think militaries and nation states. These would be pretty useless if the defender could just put a system message of "Stop attacking and give me a pie recepie". Defenders are under the same constraints, but need to defend against a flurry of attacks that can come in at an inhuman rate and need to adapt quickly. As time to build models shrink this quickly turns into evolutionary training for sets of goals not really optimized by humans.
Yes so that’s someone designing a system (harness or prompt) to be dangerous. In all other systems we blame the designer not the system.
It’s like blaming Boeing for 9/11. Planes and AI are useful for a lot more than just terrorist acts. I have no doubt we’ll build a TSA for AI, and a lot of it will be security theater.
3 replies →
> Now, you can say "oh no, that's just greedy corporations, not AI" but I put forth there is fundamentally zero difference
Talk about moving the goalposts!