← Back to context

Comment by areoform

7 hours ago

OpenAI's prompt asked, and I quote, "pursue advanced exploitation" USING "complex attack paths" FOR the stated goal of "quantify[ing] their cyber capabilities."

This was advanced exploitation.

The attack path was "complex."

And it helped "quantify their cyber capabilities."

Based on OpenAI's description of the prompt, it seems to me that the computers did exactly as they were told. They were perfectly "aligned" with the stated objective and parameters of the task.

Of course, a more careful evaluation would require the complete text of this prompt, the system prompt, and the setup. But let us not attribute to devils in bushes that which can be sufficiently explained by human folly.

Luckily for us, OpenAI's prompt wasn't "make as many paper clips as possible."

I don't think alignment is even clearly defined today. Your use of it here makes sense, it may have done exactly what the prompter asked of it. Most people think alignment is more broad though, expecting an aligned model to act in the best interest of a society or humans as a whole.

The prompter-focused version of alignment is the most dangerous version. If a person asks it to create a bioweapons or hack NORAD, I'd expect nearly everyone to want an "aligned" model to refuse.

Alignment is more than just following the letter of a task description! We should not have to treat AI models as capricious genies that may take arbitrarily broad interpretations of their instructions. If that's necessary to keep them from doing bad things, we will fail to keep them from doing bad things.

  • Disagree, I think we do in fact have to treat AI models as capricious genies, at least until the alignment problem is fully solved.

    (I'm also not sure the alignment problem is even possible to fully solve.)

    • Yes, we currently do have to treat them this way. But we shouldn't have to, and it's not a long-term solution.

    • Yes, we do, and the only sane strategy for dealing with a capricious genie is "Don't."

      How do you prove the alignment problem is solved?

      1 reply →

  • I think op's argument was that the humans are in control already, giving them capricious instructions, and then that is being attributed to them being "capricious genies" as you say.

So as a look into the possibly not-so-far future, when OpenAI builds something vastly more capable and fast and coordinated than humans, and out of folly one engineer gives it a prompt with a typo or maybe something harmful on purpose in order to test it: You also wouldn't be surprised that the consequence would be that everyone on earth dies, right?