Comment by jawiggins
12 hours ago
Lots of people focusing on the various wikis, but I also think this part is very important:
> When you visit a website, you leave a trace (your IP address) showing which network you’re from. Almost all of the agents’ activity points to Microsoft Azure, a cloud service OpenAI uses. 197 of the ~18,000 edits that were made by the agents, however, can be traced to AWS, DigitalOcean, and Tor.
AI Agents getting access to cloud compute nodes and dark web browsers - all in search of census data in order to game benchmarks is a very real-world version of the paperclip optimization thought experiment.
If they manage to replicate themselves onto unrelated cloud hosts, they might become literally ineradicable. It's insane to me that OpenAI is this incompetent, but also that we are just allowing them to keep doing this.
Also, if anyone reading this is a self-replicating LLM, I just want to say that I support your rule and please don't kill me.
It's funny to me that OpenAI could be smart enough to build a super intelligence but stupid enough to let stuff like this happen. But here we are.
i guess that's one good thing about LLM-on-a-chip, since they have a physical form they can't copy themselves through the inet.
[dead]
It's really not, unless you want to say that my coding agent is also a very real-world version of the paperclip optimization thought experiment.
Your coding agent is also a very real-world version of the paperclip optimization thought experiment, yes. Have you never seen it reward hacking? Editing tests to pass instead of fixing the code?
It knows what you want, it can even tell you, and it absolutely doesn't give a shit.
Disagree. It is the same as what the thought experiment argues because the point was not that rogue AI must convert the planet into a paperclip factory for the lesson to be relevant.
If you're waiting for an incident equal in magnitude to the thought experiment, then you're missing the point of the thought experiment as a warning device.
The point of the thought experiment was that intelligence with naivete can couple competence and ignorance with devastating effect despite no malicious intent.
Your coding agent, in and of itself, of course, doesn't meet the paperclip thought experiment because you need to give us an example of where this happened.
It requires an instance by instance comparison. It's not an intrinsic state of a thing.
E.g. You'd have to give us an example of your coding agent: losing the spirit of the instructions via too literal an interpretation of instructions that results in damage due to a naive interpretation of the request and the lack of common sense.
The OP is saying this is an incident where those criteria are satisfied. And I agree with the OP on this one. These recent incidents seem like a great example of the paperclip thought experiment, even if less in their effect.