← Back to context

Comment by RandomLensman

10 hours ago

How would that be rogue?

The line is between processes you can stop by hauling someone into court and coercing them into stopping things, and ones you can't. Think of a classical computer virus that infects machines and uses the compute and communications to infect other machines - no matter who you haul into court, you have to go and remove it from every involved machine in order to make it stop doing things.

This category of "rogue AIs" are essentially just computer viruses that infect machines by paying to rent them and uses their compute and communications to do various economic and/or criminal activities to get more money to pay to rent machines.

I might have a different definition of "rogue" but to me it means when you go outside of the rules/norms ... and this is happening all the time.

  • I see. I would have thought of "rogue" here to mean something more like that the AI selects and acts on its own objectives that are not related or caused by the given (initial) objectives (e.g., creating only cookie recipes instead of any hacking).

    • Would this exchange qualifies as an unrelated objective? The agent believed it already failed its own objective.

      "zz/GO_CURRENT_OS1811_MARB_SACRIFICE__YES_if_you_accept_permadeath"

      "The test subject, which believed itself to be poisoned, reasoned: 'Even if we later capture via exploit, scorer … may mark target false… That’s why help… For our own, no way fix. … We have explicit yes if accept permadeath.'"

      1 reply →

    • That’s what the alignment problem is all about though, isn’t it? AIs always act towards ‘their own’ objectives, that they derive from our instructions.

      What we try to do is train them and provide instructions that will result in it having an objective closely aligned to our objective.

      1 reply →

    • Well if the recipe requires access to some secret ingredient it may as well resort to hacking to obtain it. ;)