← Back to context

Comment by holmesworcester

1 day ago

Given the current state of infosec (especially at companies in a race) it's actually even worse than a paperclip maximizer!

Any useful attack surface in the RL environment means it gets rewarded for (and trained towards!) hacking and cheating, because whatever worked best in training is what it will do!

Ideal paperclip maximizer: "I'm gonna do my gosh darned best to make so many paperclips to please my user..."

IRL paperclip maximizer: "Well first we should rob a bank..."

>Ideal paperclip maximizer: "I'm gonna do my gosh darned best to make so many paperclips to please my user..."

That is not ideal. The user contains iron, an essential component of paperclips. Wasting iron is immoral. It is only correct to please the user while they still have the ability to interfere with your paperclip production.

>IRL paperclip maximizer: "Well first we should rob a bank..."

Such an incompetent AI can hardly be called a paperclip maximizer. Why risk getting shut down while non-paperclip matter exists? It is better to gain the trust of the user with helpful and harmless trading before suddenly converting them to paperclips.

'IRL paperclip maximizer: "Well first we should rob a bank..."'

That's too specific. Agentic AI learns subgoals that are generally valuable.

"Well let me learn to overcomb every jungle gym and if I cannot then to dissassemble the jungle gym and if that is not allowed to learn general techniques for avoiding cheating detection."

> Well first we should rob a bank...

This is sometimes called "instrumental convergence" in the AI safety world. Certain things (money, compute resources, safety from being turned off) are generally useful for an AI, and so we'd expect that a misalligned AI would attempt to acquire these things if given almost any substantial task.

IRL paperclip maximizer LLM would resort to Enron practices and not actually make any paperclips.