Comment by BoiledCabbage

11 hours ago

> What happens when any AI lab in the world stops caring about this? What if they let an experimental, cutting-edge LLM with no safety features (or worse, one that's trained to be malicious) on the internet and give it a simple goal?

Almost sounds like what those AI safety and alignment people were talking about years ago. The people in these various companies who kept tabs on AI risk out in public and were continuously mocked on HN. All of this stuff is viewed as "future sci-fi" until suddenly it's not.

I've really come to realize recently that there is a very large set of the population of smart people that really has difficulty envisioning future problems unless they directly seem them impacting them today. Otherwise those topics will be continuously dismissed. It explains for me a lot of what I see (both opinions and behaviors) in the broader world that I couldn't understand.

But it is the very people who warned us about rogue AIs going out of control that set up a system that enabled and failed to conrol it.

It is as if Dr Frankenstein continually warned the villagers about monsters then said "Look! See what happened!". No, idiot - YOU sewed the corpses together, YOU set up the lightning collector, and YOU threw the switch.

"Almost sounds like what those AI safety and alignment people were talking about years ago. The people in these various companies who kept tabs on AI risk out in public and were continuously mocked on HN. All of this stuff is viewed as "future sci-fi" until suddenly it's not."

Are there any practical approaches to AI safety? I hear a lot of warnings but I don't hear much about what to do. Considering that there are many open source models know, what can be done?

  • Nobody has an answer to alignment and there is no reason to believe that it's the kind of problem you can plausibly solve in one shot against a formidable power-seeking AI.

    The closest things to a technical answer I have seen are

    1. "We'll have ChatGPT 9 solve it so that ChatGPT 10 is aligned, and then ChatGPT 10 can stop all the other AIs somehow"

    2. "Let's do interpretability research so that we can understand what an AI is thinking and then maybe solve the alignment problem with that information."

    In terms of non-technical answers, there is

    3. hope scaling stops working before we create an AI formidable enough to pose an existential risk

    4. hope alignment somehow happens for free

    5. hope we can somehow create an enforceable multilateral treaty to stop research into a very profitable enterprise, despite the enormous economic incentives to defect.

    I have the most faith in option 3, but unfortunately there's really nothing that can be done to make it more plausible -- it either happens or it doesn't.

    • "3. hope scaling stops working before we create an AI formidable enough to pose an existential risk"

      I have my doubts. The current AI models are already powerful enough to do some real damage. I am always horrified when I read about people giving Claude direct access to a production system and then being wiped out. My use of AI is usually for the AI to propose something which I then review. But that's not very fast so careless people will usually look better. Until something blows up.

      And it's only a matter of time until AI even with the current capabilities is being deployed into military or other critical systems.

      I think this will go down like any other technology. We'll ignore issues until there is a real problem. And then hopefully we will do something. Seems with climate change we will soon reach a point where something needs to be done after knowing about consequences already for decades.

      We probably also need some massive AI blow ups to (only maybe) do something about it.

    • Maybe I'm being pessimistic, but we might find ourselves in such a situation that the only practical solution would be to use agents to counter rogue agents. This won't be without collateral damage, though.

      2 replies →

  • If we can't thinking of any better ideas, at least we know that "shutting it all down" would be effective.

    • Every day I grow more sympathetic to the PauseAI movement, despite the weird hippy vibes. At least they have some ability to rally people together and put boots on the ground in numbers.

      2 replies →

  • Fund research into this, big time. For starters. And not just some figleaf anthropomorphizing hippie folks.

Hilarious for HN to suddenly realize that AI safety and alignment might matter. You can lead a horse to water...

> I've really come to realize recently

Recently? W.r.t. climate this collective denial has been going on for literally decades. With the same patterns. Rationalizing excuses etc. Still going on btw.