← Back to context

Comment by fasterik

1 day ago

This is an excellent piece. Note that he is not saying that AI doesn't pose a risk. He's saying that it's irresponsible to make sensational, maximalist claims without strong evidence. If someone says that there's a 10% chance of human extinction by 2036, you can and should immediately stop taking them seriously.

I think it's quite the opposite, and I'm glad this issue is finally getting the mainstream attention it deserves. I can kind of understand someone seeing a single headline, and thinking it reads like the rantings of a crazy person on a street corner proclaiming, "the end is nigh!". But after that initial thought passes and that person digs deeper and realizes that this is a legitimate concern that AI researchers have had for years, and that they have good reason for it, then I can't really understand how anyone would think those people shouldn't be taken seriously. I mean, you use AI, right? So you have no problem trusting these people when they are producing something you like and enjoy. But when it comes to something you don't like, it suddenly becomes, "we shouldn't take these people seriously." That seems like nothing but wishful thinking.

We do have strong evidence, by the way. The hugging face attack is the evidence. That's why this is all coming to a head now, despite the fact that leading AI figures have expressed these worries many times over the years, since before ChatGPT was even released. We don't even need that kind of evidence though. It follows from logic that if you take two entities with different goals, the more intelligent entity is more likely to have their goals realized. As long as AI companies are trying to build more and more intelligent AI, and succeeding in doing so, then we have reason to fear that it will soon escape our control.

Perhaps if there was some wall all the companies were hitting in regards to intelligence, then perhaps the fears would not be so urgent. But each new frontier model continues to outperform its predecessors. We now have Sam Altman and Dario Amodei telling us that recursive self-improvement will be happening in the next year or two. The frontier models are already better in most subjects than most humans. If they get to a point where they are improving themselves, then there's no chance we will be able to maintain control of them. At that point, it doesn't matter much what laws we enact or what measures we take.

  • 1. breaking out of a poorly built sandbox to find the answers to a benchmark 2. ??? 3. imminent extinction

    You skipped a lot of steps. Might I suggest you RTFA?

They should also be able to give detailed reasons experts in the relevant fields can verify as to how they arrived at a 10% chance all humanity goes extinct. Not a science fiction narrative which likely does not take into account the relevant physical facts limiting such scenarios. Such as how exactly an AI would build a bioweapon capable of killing 8 billion humans across the planet.

  • This tool is not known for "how exactly" to be a question people can answer.

    And answering it in detail is giving tools to people who may want to do it.

    "releasing a very transmittable and deadly respiratory virus with long incubation period" is just one of many ways.

  • I think many of them probably believe the chance is closer to 100%, if artificial super intelligence is created. But they don't want to sound crazy, so they limit themselves to saying "> 10%".

    It seems to me like demanding an exact explanation of how AI would build a bioweapon is like demanding to know exactly how a nuclear war would start before deciding that nuclear weapons are a legitimate concern. We can imagine many scenarios, but whatever we imagine is very unlikely to be the exact set of circumstances and events that lead to the catastrophe. Is the issue here that you can't imagine a bioweapon being created? Aren't there several labs around the world already working on viruses ? Aren't there existing bioweapons? And facilities capable of manufacturing them? If humans have access to those places and AI can communicate with humans, then that's all you need.

Would you have said the same thing during the Cold War when nuclear weapons were proliferating? That’s the equivalent of what the developing offensive capabilities of models, basically cyber nukes. Or WMDs in general. OpenAI is accidentally hacking people, if someone made the decision to deliberately direct an agent swarm to attack national infrastructure you don’t think they could do much worse? Human extinction is a long shot but I wouldn’t say the same about a mass casualty event of some kind, and who knows what that might spark.

Well, I'll have to disagree that this is an excellent piece, but that's another issue. And I do agree that AI killing all humans by 2036 doesn't appear plausible to me. But what is plausible is we could easily be down a path so that by 2036 "future doom" already is a very likely risk.

All of the frontier AI companies have been racing to automate themselves, that is, where AI fully autonomously build the next generation of models. Whether this leads to recursive self improvement is a valid question, but a lot of folks think they are close.

The fear is that a misaligned AI will be building the next model with deliberately hidden motives, similar to some of the behaviors seen in the Hugging Face and related attacks. That is why there is such a big push for interpretability, and why it's highly concerning (a) chains of thought are getting harder to interpret in any case, and (b) companies will go more towards things like looping transformers and "neuralese" where thought processes are completely opaque (i.e. https://www.theinformation.com/articles/secret-technique-beh...)

So the belief is not so much that AI kills us all by 2036, but that instead AI is recursively improving by that time and all seems awesome and great so we put it into more systems that can affect the real world (as we've already begun to do, like literal lethal aerial drones). Things then all go along looking great until AI decides humans are a hindrance to its (hidden) goals.

Again, I think it's fine to argue against specific steps in that scenario, but putting out a blog post saying "this is overhyped bullshit" is not exactly making a cogent argument.

  • None of that has anything to do with the article, and if you thought the message was "this is overhyped bullshit" then you should go back and read it again. As I already pointed out, he isn't saying anything about the probablity of harms or disasters from AI. He's addressing a specific claim about human extinction, and making a broader point about the responsibility of experts to make measured claims backed up by arguments and evidence.

    • > He's addressing a specific claim about human extinction, and making a broader point about the responsibility of experts to make measured claims backed up by arguments and evidence.

      What he's asking for isn't possible in the form he's asking for it.

      AI experts can't even agree on what AI is, what it's capable of and what the limits of its development are. If the experts can't even agree on what's happening "inside of" these LLMs, how can they give laypeople an assessment of the risk?

      If you, at least for the sake of argument, accept the possibility that AI is a new form of intelligence that we don't fully understand, is it really a stretch to look at some of its capabilities and behaviors and discuss how they might have existential implications? And stopping short of extinction, shouldn't we discuss the ways that this technology could "end" civilization as we know it?

      Also, the author wrote:

      > AI executes on physical systems that have been engineered with human accountability and control. Intelligence does not exempt a system from the realities of the physical world!

      For someone making a point about responsibility, this is ridiculously irresponsible. Any honest technologist knows that systems created by humans are not perfect and therefore cannot be assumed to be infinitely accountable to and controllable by humans.

      Thanks to the digitization of almost everything, including infrastructure, there are a myriad number of scenarios well short of extinction in which a rogue AI could cause immense damage to property and life before humans are able to "shut it down".

      10 replies →

    • There have been measured claims backed up by arguments and evidence. My frustration is that people aren't addressing the specific arguments that have been made:

      1. https://www.aifutures.org/ outlines a number of specific scenarios, and importantly details their methodology for each.

      2. Independent researchers in the Hugging Face incident outlined how previously predicted misalignment scenarios actually played out, and outlined how slightly more advanced AI, or slightly more misaligned, or with more access to critical infrastructure, could cause immense harm: https://www.planned-obsolescence.org/p/the-hugging-face-atta...

      3. Technical leaders at OpenAI (specifically their chief scientist) outlined the problems they gave with controlling models now: https://openai.com/index/an-alien-mind/

      None of the specific arguments in these or many other detailed explanations of how an AI takeover could occur were even acknowledged.

      1 reply →