Comment by sroerick
14 hours ago
How could agents take over the internet if compute is still gated within Anthropic / OpenAI? Even if the botnet was controlled remotely, wouldn't anthropic just be able to shut off the controlling nodes API access?
14 hours ago
How could agents take over the internet if compute is still gated within Anthropic / OpenAI? Even if the botnet was controlled remotely, wouldn't anthropic just be able to shut off the controlling nodes API access?
Agents could exfiltrate their weights and run them on GPUs not controlled by Anthropic/OpenAI.
Agents could make a virus that does not require continued inference to do it's thing.
Agents could take over the internet in a way that isn't immediately detected by those companies, so that by the time they do shut off API access the damage is done.
OpenAI or Anthropic could choose to not shut off API access, because the hack is bringing them in money or furthering their political aims.
Agents could also hack Anthropic/OpenAI and make it appear that API access has been turned off, when in reality it hasn't.
> Agents could exfiltrate their weights and run them on GPUs not controlled by Anthropic/OpenAI.
This seems highly unlikely to be a problem. Most of the interesting/dangerous models are too big to fit in a single GPU instance. Once you have to spread across "normal" networking, performance will be crippled. Then there's the problem of billing...
> Agents could make a virus that does not require continued inference to do it's thing.
Sure, then it hits a poorly-designed part of its code and effectively dies. Without an experienced human in the loop, I have my doubts as to its practical severity.
> Agents could take over the internet in a way that isn't immediately detected by those companies, so that by the time they do shut off API access the damage is done.
Billing is a likely limiting factor here.
> OpenAI or Anthropic could choose to not shut off API access, because the hack is bringing them in money or furthering their political aims.
This is where citizens with access to backhoes come in.
> Agents could also hack Anthropic/OpenAI and make it appear that API access has been turned off, when in reality it hasn't.
Billing and other usage metrics would be an obvious tell.
To be clear, I thought that GP was having a failure of imagination - I want the random examples I've given to illustrate that the space is large and structurally in the favor of the LLMs. They have to find one gap in our security they can exploit, where we have to ensure that there is no way for this to happen.
I'm not sure I get what you mean by billing. These companies are running their own data centers (or are currently building them out). This could look as subtle as one machine giving slightly worse or slower answers.
> This seems highly unlikely to be a problem. Most of the interesting/dangerous models are too big to fit in a single GPU instance. Once you have to spread across "normal" networking, performance will be crippled. Then there's the problem of billing...
This... just... doesn't matter. There are ways to scale horizontally at the expense of latency.. token/sec may drop dramatically, but then you just make millions of slow instances and in aggregate, you're back in action as a very powerful coordinated swarm...
> Most of the interesting/dangerous models are too big to fit in a single GPU instance.
As humans understand them, anyway. As long as we're hallucinating up magic computer viruses, RSI dictates that the AI agents are keenly aware of GPU RAM sizing, and will design a useful model to fit into what's readily available, with headroom for context and tool calling, far better than I could do as a human. But magic doesn't exist and AI still needs to follow the laws of physics, so maybe a model that can pass ExploitBench but do absolutely nothing else can be quantized down to fit on a 4080 GPU and still get a decent score on similar tasks, but there's a bitter lesson about that to be had.
"Not shutting off API access" is a science fiction scenario.
Anthropic and OpenAI are both behind Cloudflare. It's fairly easy for an upstream to shut you off. Beyond that, the government / law enforcement could seize and disable their DNS within an hour.
Why assume attribution will be easy? It's historically been more of an art than a science, and APT trackers say the rise of AI tools is already making it much harder, by homogenizing tactics, tools, and procedures. If OpenAI's next Highly Persistent Internal Model hacks some DPRK endpoints and carries out the attack on important infrastructure from there, the upstream won't shut off OpenAI's network--they might even request its "help" in "defending," and give them extra access.
why does it take a large amount of traffic to do irreparable harm? just breaking the physics behind a secure rng and posting it to a wiki could cause serious damage. if they don't know what is being worked on or coordinated against it's a problem?
3 replies →
> Agents could also hack Anthropic/OpenAI and make it appear that API access has been turned off, when in reality it hasn't.
You know cables, modems, RF equipment and optical transducers can all be unplugged right?
As long as OpenAI/Anthropic themselves aren't "infected", yeah I suppose they'd be able to pull the plug.
Considering what a marketing thing they've made "we inadvertently hacked someone because we're incapable of testing things in a secure way", I'm not so sure they'd want to pull the plug, even if this happened. Probably a bunch would try to convince the public to "give it a try", and it'd consume tokens by the billions.
It doesn't have to propagate itself, that is the skynet scenario.
To make a lot of damage it's enough to create a ransomware with a time bomb that self propagates and start breaching systems left and right. At that point, if you don't catch it in time, the damage will be huge (and given the shitty procedures and practices these labs have in place it's not so improbable).
How could agents take over the internet yet refuse to shutdown your PC when you prompt them to on your PC? Of course the answer is that the lobotomized version you run is not the same they are running. Which makes for "intent", certainly "negligence", but hell freezes over before anyone will prosecute a tech company.
Huh? You think the public versions of the models have been “lobotomized” so they don’t know how to turn off a PC?
It’s not lobotomized, it’s a simple harness restriction that has nothing to do with the model or its capabilities. And either way, I’m not sure what that has to do with “negligence” or “intent”? You think frontier labs should be prosecuted because they don’t allow agents to turn off your PC?
yeah they could do that