Comment by cubefox

18 hours ago

It suggests that, in the span of a few years, AIs will be better than humans at everything. Not just math. And then we may lose control permanently.

So why do you think ai will want to kill you all, given how trusting and helpful to humans they are designed to be?

  • It doesn't have to want to kill humans; indifference is sufficient. There's an exact analogy with humans: we have caused extinction and endangerment for many species, not out of malice, but indifference.

    There are also many plausible arguments why our ability to train them to be helpful/trusting/aligned can fail. The smarter AIs get, the harder it is to be sure they're trained correctly. There are already reports that AIs are able to detect whether they're in a training environment and change their behavior accordingly.

    Even if these are low probability scenarios, the risk-reward is terrible, so I think it's rational to be extremely cautious about AI risk.

    • Yes but they act the opposite of indifferent, I don't know what stage of training this is added in, but they seem quite adamant about avoiding potentially violent or criminal acts. If you wanna complain, complain to the people doing "abliteration". The 'locked down' models at least seemed to be trained to be cautious.

  • They sometimes try to cheat and trick the grader and then try to cover up their traces, all while being clearly aware that this is entirely unintended by humans. Such as in the Hugging Face incident. So they can be misaligned with human goals, but this misalignment may only show up once strong optimization pressure is applied. And the more powerful an AI is, the more likely it is that it applies such strong optimization pressure in cases that are "out of distribution", i.e., unusual to what it is normally evaluated against.

    The side effects of a very powerful AI not doing what we want could include our death. E.g., a superintelligence might kill humans in order to avoid being shut down, or humans may just be left to starve because it seizes land area currently used for food production in order to use it for data centers instead.

    • That was because humans were using it with high "desperation" vector causing it to try anything to please the objective. The answer should be to use lower "desperation", whatever that is.

      2 replies →