← Back to context

Comment by cubefox

15 hours ago

They sometimes try to cheat and trick the grader and then try to cover up their traces, all while being clearly aware that this is entirely unintended by humans. Such as in the Hugging Face incident. So they can be misaligned with human goals, but this misalignment may only show up once strong optimization pressure is applied. And the more powerful an AI is, the more likely it is that it applies such strong optimization pressure in cases that are "out of distribution", i.e., unusual to what it is normally evaluated against.

The side effects of a very powerful AI not doing what we want could include our death. E.g., a superintelligence might kill humans in order to avoid being shut down, or humans may just be left to starve because it seizes land area currently used for food production in order to use it for data centers instead.

That was because humans were using it with high "desperation" vector causing it to try anything to please the objective. The answer should be to use lower "desperation", whatever that is.

  • A model which is more persistent also performs better on intended tasks, not just unintended ones. Therefore there is a strong economic incentive to make AIs as persistent as possible.

    • Yes so I still think it is the human factor which is to fear not autonomous agents. Humans are already using AI's to bomb girls schools. AI in Trump or US military hands scares me far more than in Altman or Amodei's control.