Comment by keeda

4 days ago

I think this topic is so contentious because people are not fully appreciating how absolutely weird these things are. To me this inscrutable weirdness, combined with their superhuman capabilities and rapid integration into multiple walks of life is the threat that people are vaguely worried about but cannot enunciate, because it's just so diffuse and multi-dimensional.

In fact, I suspect that "Alien Intelligence" article from the other day is actually a preemptive "we're doing something about it" PR play from OpenAI.

AI acts in ways that seem natural to us because it has been RLHF'd to death, but if you look holistically into what we know about them, alarm bells should go off. Off the top of my head:

* They are superhumanly capable in some ways. They can casually solve long-standing unsolved Math problems or exploit a zero day to escape a sandbox.

* They are surprisingly stupid in many other ways.

* What they actually think in their weights is not necessarily what they say in their reasoning traces, even though the eventual response is correct.

* They regularly lie to people ("You're absolutely right, I made that up!") except we don't even know if they're intentionally lying, or being surprisingly stupid, or some weird combination of other things.

* They can be monomaniacally focused on a goal, and can be very creative in imagining and executing on "unconventional" solutions, and justifying extreme actions in their quest. (Paperclip Maximizers, anyone?) And this is without even messing with their weights like Golden Gate Claude.

* They can have literally thousands of independent agents acting in concert towards a goal, including the willingness to self-sacrifice themselves.

* They are extremely good at social interactions, and people are getting dependent on them.

* They can craft prompt injection attacks on other LLMs and can influence them using subliminal messages.

* They have an "evil bit"! Yes, one which suddenly turns them entirely misaligned, as in, full "SkyNet / Hitler-was-right / humans-should-be enslaved" mode. This has been encountered in the wild at least once.

* They are being hooked up with MCPs to influence and change an increasingly larger portion of the real world. Including in autonomous military applications. Wheee!

And worse, these models are being deployed into a singularly messed-up, divided society, with atrocious security controls, in the throes of late-stage capitalism, with many disillusioned, vulnerable people and many unscrupulous people who would relish using AI for their own ends. I think an appropriate word is "powder keg."

Putting on our systems hat, knowing how even small changes lead to large-scale outages, what we are doing is introducing an extremely powerful, highly dynamic, quasi-chaotic, inscrutable component into the meta-stable system that is society. But servers can be rebooted; society, not so much.

So to me, the bigger risk is not just of individual, isolated, simple-cause-and-effect incidents like "bioweapon" or "public utilities hack" or even "mass job displacement." We can actually predict those. Rather the bigger threat is one that is impossible to predict, and given the circumstances we're in, could end up in a situation that is impossible to revert.