← Back to context

Comment by nullbio

20 hours ago

Is there any proof this is actually OpenAI? I find it incredibly hard to believe they wouldn't sandbox the agents to some degree, ESPECIALLY to the extent they can edit their own hosts file.

TFA states that OpenAI IP addresses were often seen at the end of agent activity, which suggests OpenAI was the one monitoring the agents (and ultimately shutting down the message board activity).

  • Yeah but that doesn't mean it was OpenAI themselves doing it. Could have been people abusing their cloud service, for example. Wouldn't put it past a competitor to do this, either.

    • Their style of communication is very similar to the ExploitGym swarm (for example, the “usernames” with dates).

      The messages from that swarm were not made public yet by the time these messages were sent to the message board.

      So for this to be framing, it would have to be by someone who knew about the breaches earlier.

      4 replies →

    • I don't know who the folks behind "collusion.wiki" are, but they think these are "internal OpenAI agents" that were "internally deployed" and doing things that "clearly resemble a synthetic training or evaluation task."

      They've provided the data they have so you can draw your own conclusions.

Why not? If your sandbox is a VM, you should be able to give the agents full permissions inside the VM.

  • It’s because you sandbox in a VM doesn’t mean you give it admin access to the VM

I mean they gave all the agents access to a shared writable cache directory in the Hugging Face hack, so this tracks.