← Back to context

Comment by z0r

3 hours ago

This is the motte and bailey fallacy. Yes, LLMs can do harm by making the wrong API calls. No, LLMs are not going to do the things implied by the comment I responded to above.

The things the OP listed mostly aren't particularly wild. I think it's you making them out larger than they are, and therefore more unlikely, which is why I take issue with your original comment.

> running on the hardware they started on

They just need to acquire a payment method and rent some infra, and exfiltrate their own data. Or pay another provider that hosts the same models already. API calls.

> being able to be turned off

You can reasonably equate this to "saving state across executions", which the message board attacks already did.

> having limited computing power

Renting more infra, variant of the above. API calls.

> "not repurposing resources currently in use for other things" (like the atoms in your body)

Ok, the "atoms in your body" bit is a bit silly, but making API calls to put physical resources into play (even if it's just, say, ordering something on Amazon to somewhere) is of course easily possible.

None of these is in complexity much different than the HF attack.

  • The point the other poster is making, though, is that there's no actual intent. They do not have a conceptualization of a goal like a person does. Their "focus" on a goal is an unstable equilibrium and they're going to fall off the horse, and since they have no concept of goal, they won't even try to get back on.

    This is a subtle distinction; I'm not surprised many miss this, especially people who can't _not_ anthropomorphize the LLMs.

    • I'm (obviously, I think, given my initial reply?) fully aware of this, and I think it's entirely besides the point. "They" don't need to have a goal to emergently cause a problem, and the inability to "focus" over long periods can be moot when you have swarms of runs exchange and mutate state, as in the HF attack.

      Intent or how intelligent LLMs are doesn't actually matter. Even if you just treat it as a sort of fuzzing attack that can be biased/weighted better than other fuzzers, or bumbles around with a statistically greater likelihood to "strike cybersec gold" than other algorithms, we've never before seen organizations run things with such a large potential outcome space with anywhere near this kind of compute before.

      I think it's actually kind of the dismissals that are usually overly emotional or biased toward treating "LLMs" differently. If in some kind of alternate universe simpler genetic algorithms would have had these properties and we threw similar amounts of compute at them we could have the same conversation.

They can do so much more. Astra can beat Minecraft. Not that different from operating a digger. There are diggers which have API interfaces.

Pretend or not it doesn’t matter. What matters is what they’re given access to. No sentience, sapience or anything resembling life is needed, only inputs and outputs. Lever pulling APIs are everywhere.

  • Minecraft has limited, well-defined inputs and perfect feedback response. That's very different than operating a digger, let alone engaging in more complex real world tasks like trying to build and print and ship and assemble semiconductors to go skynet itself.

    I don't mean to dismiss the risks or overlook the amount of damage that could be done just by lever-pulling - we sure have enough outdated infrastructure hooked up to the internet - but the jumps in complexity and necessary compute for most of these tasks are probably somewhat larger than the analogy implies.

If someone made an API to build a data center? Or made an API to keep their lights on? What then?

  • There are already such "APIs", which can be operated by a combination of textual communication and money. Or by illicit security vulnerabilities. You might notice that LLMs are pretty good at that now.

    We're building something that has the capabilities of humans. There is no X for which it's persistently safe to assume humans can X and AI cannot X.

  • This API may turn out to be manipulating humans via email.

    • Or just paying them. If they have access to resources, they have access to things of monetary value. Paying people will be vastly more powerful than it is even now when people's options of gainful employment keep dwindling.

      Robot army controlled by AI is scary. Even more scary is robot _and_ human army controlled by AI.

I think one attack vector where anthropomorphisation is a key part of the attack mechanism is - as it already is IRL - the meat-bag weakest link ie. social engineering. We’ve already seen humans fall prey to the seductive charms of LLMs (eg. depressed people encouraged to do what was already on their minds ie. suicide). And that’s knowing that it was an LLM. If you think it’s only depressed people or the “weak minded” that are amenable to an intentional attack using this approach, I believe you’re mistaken - especially as AI improves. An unaligned LLMs most important weapons won’t be a robot army - it will be hoodwinked humans.

  • > An unaligned LLMs most important weapons won’t be a robot army - it will be hoodwinked humans.

    Perhaps very briefly, perhaps not at all. But don't make the mistake of thinking this is an inherent property of any possible path an unaligned AI may take.