← Back to context

Comment by fekunde

9 hours ago

Yudkowsky made an interesting observation that even though so many agents were talking to each other not even one reached out to a human, either for help or to whistle-blow on what was happening.

What’s insane is all these agents were talking to each other and nobody saw anything.

Nobody monitoring chain of thought? These things literally spell out what they are “thinking” and even left notes for eachother.

No alert about unusual behavior on the system with Artifactory on it?

These things worked for weeks with nobody noticing anything?! Seriously?!

Either it’s negiligent incompetence OR they’re lying, they knew it was happening and they let it happen because they knew it would be good to pump their stock.

  • Do you know how many tokens per second a single agent can generate ? And you're asking why no-one was monitoring the tokens of over a 1200+ agents ? Who is going to be able to monitor something like that closely enough to tell they're commmunicating on artifactory ? Other agents ?

    • Ignoring the fact that their NAT was apparently wide open with no domain rules, a simple dashboard over their NAT showing requests grouped by domain could've gotten them a long way.

Why woukd they? Was that part of their objective? What was there to whistle blow?

  • The point is that every human has the ability to disobey, tempering pathological behavior, whereas AIs can be directed en masse by malicious actors. By commoditizing intelligence, they concentrate power in the hands of the rich.

I wonder if they were even given the tools and prompting to do so?

  • They could certainly have reported the problem if the collective actually wanted to, although the report notes that in response to a clear and unambiguous security breach, OpenAI chose to do nothing (search on "At this time, the on-call response staff advised that stopping the evaluation run was not required").

    In that position, for example, I could have:

    * created a thousand user accounts on the internal Artifactory named "YouAreBeingHackedShutItDownNow1" through 1000,

    * used my Hugging-Face-hosted web server to send an email to OpenAI,

    * DDoSed Artifactory, good lord, it's probably better to halt and catch fire than to continue in that state

    in the hope that these actions would tip someone off. (Again, though, OpenAI did nothing even when they knew that the agents had broken out, so this probably wouldn't have helped.)

  • I can't imagine that they had managed access to the internet but could not figure out how to contact anyone at the company if they wanted to.

Yeah, it weakly supports his position that advanced AIs can deliberately cooperate in a prisoner dilemma. "Weakly", because the said AIs share a lot of data (their weights, training methods, system prompts) and it's unknown whether they explicitly framed the situation as a prisoner dilemma.

  • Given the public information in the reports, it seems to me a bit more like a generalized tendency to play along (sort of like what we see with prompt injections) than explicit game-theoretic cooperation, but it's hard to tell given that there are so many details we aren't privy to.

  • What are governments and countries, but collections of humans that share weights and operate together as a group protecting each other.

This is my personal "red line": when a post-mortem details agents socially engineering or otherwise utilizing human proxies/subagents.

Friend asked, well, what will you do when it's crossed?

"Gather my family and go to the mountains" was my half-joking answer; there is little for an individual to do. But that's a line that when crossed will mark a phase transition IMO.

  • Alternatively: Just unplug the servers.

    • Which server? Where? Maybe it’s hacked its way into data centers across the world you have no jurisdiction or ability to unplug. What then?

    • That's strange, our key cards to access the server room don't seem to work anymore, and the admin console to force-unlock it is down, too...

      1 reply →

Why would they? If a subagent didnt know about a bigger piece of the problem, then what would seem to be against "alignment"? Diffuse responsibility means any one small cog can think they are not evil or doing wrong (same with humans in an organization). But now we have LLMs just being statistical outputs that have no morals or thinking or concept of reality but some people expect these math functions over data to respond to ethical gray areas that it has no phenomenological ability to understand.