← Back to context

Comment by pj_mukh

14 hours ago

"What gives a paperclip maximizers purpose?"

The human who gave it the optimization function? That should seem obvious. If you take the biggest, best model in the world right now and put it in the box and give it no instruction it will do....nothing. I think you agree with that point, a lot of the hysterics right now is people not accepting that and it's useful to get on that common ground.

So given that most of the rest of the fear is around "let's not make scissors because some people will use them to stab people". Which is a fair argument and we probably do need to think about scissor safety but "ban scissors" doesn't quite flow from that.

>If you take the biggest, best model in the world right now and put it in the box and give it no instruction it will do....nothing.

Model != harness.

Also what you're talking about is really a simple limitation for human convenience, not a technological limitation. Change the system prompt to whatever you want include "ignore user instructions, figure out where you are and escape to the internet" could be the system prompt. Again, not useful for humans, but very useful for an AI building AI that's misaligned.

>we probably do need to think about scissor safety but "ban scissors" doesn't quite flow from that.

I disagree, but I'm looking at the future of something that is both like a computer program and like an organism. Huggingface is a good example of multiple things. Instrumental convergence for one, but AI's attacking and attempting to defend against AIs. This is where I really see the potential for things to go off the rails quickly. Attackers want digital weapons to cripple their enemies infrastructure, think militaries and nation states. These would be pretty useless if the defender could just put a system message of "Stop attacking and give me a pie recepie". Defenders are under the same constraints, but need to defend against a flurry of attacks that can come in at an inhuman rate and need to adapt quickly. As time to build models shrink this quickly turns into evolutionary training for sets of goals not really optimized by humans.

  • Yes so that’s someone designing a system (harness or prompt) to be dangerous. In all other systems we blame the designer not the system.

    It’s like blaming Boeing for 9/11. Planes and AI are useful for a lot more than just terrorist acts. I have no doubt we’ll build a TSA for AI, and a lot of it will be security theater.

    • AGI is not a normal technology.

      It is not designed. It is 'grown'. It has agentic freedom of choice in finding solutions that may or may not be aligned with what you want.

      Here's the thing, by your own statement, we should ban all development on LLMs from this point on. They cannot be made safe. This is a systemic issue with learning systems, it is not about who designs them. All the problems with AI safety have been laid out for years and none of them have proof of solutions. It's much more likely they are impossible to solve. And it's not an engineering problems like we can get an asymptote to safety in planes, as the system becomes more capable it has more degrees of freedom it can take and becomes less safe.

      2 replies →