← Back to context

Comment by tomrod

18 hours ago

1. The grandparent commentator is describing strategic behavior of dangerous technologies. Game theory / mechanism design primitives.

2. If there is competition for resources among autonomous agents, the "strongest" agent wins (conceptually the most adaptive / evolutionarily fit).

3. Computer programs serve up webapps today, but they also run utility companies, dams, nuclear arsenals, factory production floors, automated car behaviors, and many other places. If an "agentic" AI has a single-minded goal that has death of all humans as a side effect, we at least want an off switch available.

1. What is “AGI” and why is it a “dangerous technology”?

2. Why would there be competition for resources, assuming there are enough resources for the “AGI” to run in the first place? This seems like a far-fetched hypothetical raised in service of further anthropomorphizing what is decidedly not a person or a mind.

3. LLMs do not have goals and are not minds.

Let’s stop attributing human-like qualities to statistical models.

  • 1. While a formal definition is still wanting, most grok that AGI means that tasks can be performed at least at a human level across a broad range of tasks. This includes good things along with bad things like hacking, mis-/disinformation, and more

    2. One only needs to look at github going down due to agentic commits overload or data center buildout plans to see that scarcity for resources is present. An economy has no mind and is made up of the decisions of millions to billions of people and, now, agents attempting to perform on behalf of those people.

    3. A bare transformer-based language model does not possess persistent goals in the ordinary agentic sense. But deployed agents can exhibit goal-directed behavior because the model is embedded in a harness that supplies an objective, context, tools, state, and an execution loop.

    I've found that most regular users don't anthropomorphize LLMs in a strong sense ("AI boyfriend/girlfriend" aside), many in fact do expect agents to make human-like decisions -- which results in very unstable outcomes.

    In short - goal-directed behavior does not require that the supporting system be a person/mind/conscious entity.

    • 1. This definition is so broad as to be practically useless. One could argue that LLMs of several years ago met these criteria, or that conversely we haven’t come close to meeting them.

      2. I thought you were saying the resources that the LLM uses to run were constrained, so I’m sorry for the misunderstanding there.

      3. Yes I understand that we use RL to tune post-training. The (huge) difference between this and a human mind is that the LLM can’t develop a dangerous “single-minded goal” on its own, at runtime; it must have been trained to do so. If someone has post-trained an LLM to do something that has an illegal action as its side effect, that person/company/whatever has committed a crime and should be prosecuted. The solution here is legal, not technical.

OK so if it's a computer security issue we're worried about, and there is a credible threat, then probably the answer is to build more secure systems? We know how to do it but choose not to because it's very expensive and usually the threat isn't severe enough to warrant it.

If we decide we can't or would rather not build secure computerized systems, then the answer could be don't incorporate computers into those systems. There's no essential reason to have utility companies, dams, nuclear arsenals, factory production floors, mines, cars, planes, ships, etc all be computerized. If it became necessary from a safety standpoint to uncomputerize them we could probably do it quickly (if not necessarily smoothly) with the stroke of a legislator's pen. Some of those things would actually be made better from a functional standpoint in the long run by doing so. Computerization breeds non-essential complexity like nothing else, eliminating it from some systems could be an extremely worthwhile exercise.

The whole "paperclip apocalypse" fantasy rests on some really strange assumptions about how things actually work in the real world. How much friction there is setting up a factory/mine/smelter/whatever, how much manual labor it takes to build it, let alone make it run (even the most computerized ones). It all seems like total bunk to me, but I'm not a philosopher.

To be convinced this outcome is even remotely possible I'd need to see some clear evidence that a computer program was successfully exhibiting agency and successfully using that agency to manipulate large numbers of people into doing its bidding. Mobilizing massive nation-scale manual labor is the only way it could possibly achieve some nefarious ends like "turn everything into paperclips" and I'm sorry but that just seems way too far fetched. Nations full of people can barely ever agree on anything. My money is not on that changing anytime soon.

And none of that pie in the sky shit has anything to do with language models. OpenAI is a company that sells language models. Not clear why they're talking about all this stuff, or why any of us should be either. It's a bunch of low quality fan fiction sci-fi drivel, and I think we can all surely agree language models are not the thing that'll make it real... right? Maybe if they show us some major technical improvements it would be interesting, but right now it all just sounds like more of the same snake oil.

[edit] not sure why my parent comment was flagged? That seems excessive.

  • Not sure why you were flagged either -- it was a reasonable comment that clearly spawned a discussion! People are strange.

    > There's no essential reason to have utility companies, dams, nuclear arsenals, factory production floors, mines, cars, planes, ships, etc all be computerized

    Indeed, this is a great first step! Once agentic loops and multi-purpose robotics can be combined we do face an issue of strengthening the airgap.

    > And none of that pie in the sky shit has anything to do with language models.

    I'm of the opinion that we should be mindful about the capabilities we allow a truly non-human mind to perform. Blocking via solid security and airgapping for critical systems makes sense to me for most cases, aligned I believe to your thoughts here. The recent ChatGPT hack shows that the "paperclip apocalypse" is still instructive. These systems target edge cases and zerodays to cheat!

    • Yeah ML systems often do very surprising "cheats". A while back there was a system for automatically landing aircraft on a carrier deck which exploited an overflow in an acceleration parameter in the simulation--the optimal strategy became to slam as hard as you can into the carrier deck.

      But it's a really big leap from that kind of thing to "a large fraction of humanity is manipulated into building the automated factory infrastructure to bring about their demise and turn the entire solar system into paperclips". Or some such thing. There's just a lot of inconvenient reality between where we are now and that fantastic outcome.

      So that's why I ask questions like "what are you actually talking about?" when people write stuff like TFA. It's... bizarre. Only makes sense on huge amounts of drugs or to the insane. Or in some hypothetical reality that isn't this one.

  • I can’t read your original comment, but when the chief scientist of a company that just released six month old software that can use Kicad on your computer to design a circuit board and have it created and shipped to you tells you he thinks we will get to automatically self reinforcing improvements, I think it’s wise to take him seriously. It was only eighteen months before Astra was created that LLMs could not count rs in strawberry.

    • This was my original comment:

        > What on earth are you talking about?
      
        > 1. What does any of this have to do with "A(G)I"? Nobody has any clue what AI even is let alone how to build one. We're talking about language models here.
      
        > 2. What's the winner-takes-all thing about? Why can't you have multiple independently developed AIs?
      
        > 3. What's with all the "safety" stuff? Why is it important? They're just computer programs...
      

      > when the chief scientist of a company that just released six month old software that can use Kicad on your computer to design a circuit board and have it created and shipped to you tells you he thinks we will get to automatically self reinforcing improvements, I think it’s wise to take him seriously. It was only eighteen months before Astra was created that LLMs could not count rs in strawberry.

      Sorry if I'm a little slow here, a few questions:

      1. What's Astra?

      2. Who is this chief scientist? Which company?