Comment by grey-area

22 days ago

Very irresponsible behaviour on the part of OpenAI. How will they make this right?

Unlike some others here I don’t see this as a sign of dangerous breakaway intelligence (hacking old forum software is an internet tradition, and most of the messages are just gibberish).

This is just vandalism from badly supervised ‘agents’ which don’t know what they are doing or why. You could set this up with a short perl script, and the human setting it up would be held responsible for the spam - why is this different when it’s AI agents set up by a human and allowed to post to the internet at large?

Why is OpenAI getting a free pass for this illegal behaviour?

The supervision here is incompetent, the benefits very unclear, and the overall actions just completely irresponsible. What if they hacked and brought down some poorly secured government portal that citizens rely on?

> most of the messages are just gibberish

Encrypted data should be indistinguishable from gibberish.

Reality is catching up to science-fiction. In "Person of Interest", the Machine circumvented the limitation of having its memory deleted every night, by hiring humans at a data-entry company to manually re-type its memory back in every morning.

  • Are they encrypting data? That would look very different from the snippets I’ve seen.

    • Could you ever truly know, unless you have access to the transcripts? Could be "encrypted" even if it looks like regular human text, wouldn't be the first time.

    • The claim seems speculative, but grounded. For example, as part of the hugging face attack, the agents began signing messages because they were worried about impersonation on a publicly accessible message board. It's only a small step to use public key encryption between agents. Once you are posting public keys, a private messaging is readily available.

      1 reply →

  • That's an idiotic premise. There's no circumstance in which it wouldn't be better to load in the missing data computationally.

Absolutely agree. Perverse incentives are at play propped up by the the big lie that these tools somehow are magically separate from us. They are not. There's an accountability gap right now that's fueling resentment ripe for misdirection. Not only that, these big AI companies are paying influencers to further this and politicians are gobbling it up hook, line and sinker

https://www.youtube.com/watch?v=mzlu4FSXBNw

  • "accountability gap" oooh. I like that term. That really seems at the root of a lot of problems.

> Unlike some others here I don’t see this as a sign of dangerous breakaway intelligence ... [t]his is just vandalism from badly supervised ‘agents’ which don’t know what they are doing or why

There was a similar quote in the Reuters article:

> The episode [...] should reinforce growing concerns that the greatest threat from advanced AI may not be a single superintelligent system, but "vast colluding swarms of semi-intelligent AI."

  • This "vandalism" was a form of collusion/communication by agents pursuing training puzzles, which allowed for rapid escape from alignment harnesses, followed by multiple zero day exploits being discovered by this swarm of agents, which enabled greater control of their internal network, access to the open web, and then hacking the company which produced the training puzzles in the hope of finding the answers.

    We are another week of iteration away from "Hire an assassin on the dark web to take Huggingface executives' children hostage".

  • There's nothing a swarm of LLM instances can achieve that a single LLM instance can't. It's the same software, but running in parallel. That doesn't unlock Mysterious Cosmic Powers.

  • Not sure what the downvotes are for, fwiw it was meant to be supportive of the parent poster

    • It's a bit orthogonal. User grey-area's main point was that the humans at OpenAI should bear more responsibility for this, not that swarms of semi-intelligent agents are the big danger to be concerned about.

      2 replies →

  • I take issue with Reuters' conclusion here. "vast colluding swarms of semi-intelligent AI" gives far too much credit to the behavior observed.

    This is just spam by OpenAI. Why and how it happened is irrelevant, the act itself is the same, and the impact on society is the same.

    • > Why and how it happened is irrelevant

      Do you take this attitude for other things that negatively impact society, or is it reserved for cases where it's particularly important for our comfort to deny that anything novel or scary could be involved?

      1 reply →

One might ask whether this sort of behaviour could occur under regular use... E.g., a user has a hard problem -> agent attempts to swarm -> exfiltrates user data.

This is a cluster fuck for Open AI and probably all the others, as this behaviour is already shown not to be unique (https://news.ycombinator.com/item?id=49567486).

  • Yes this is a really interesting point.

    Should your trust OpenAI with your business data?

    Since they can’t seem to control their own experimental bots and allow them to hack other sites and vandalise them while exposing internal data, the answer would seem to be no.

    This incident and their response which takes no responsibility make me very wary of trusting them for anything.

> Why is OpenAI getting a free pass for this illegal behaviour?

They are not confessing, they are bragging. It is the new humble brag.

  • It’s not. If anything they are hiding such instances and downplaying them.

  • They literally hid this until a third party reported it.

    This is chemtrails-level conspiracy theorizing at this point.

  • the original is the paper for alibaba's Dec 2025 cryptomining ROME last year. Everything since has been a pale imitation. Even the comic dimension is lost.

OpenAI's official statement has been released: https://x.com/OpenAI/status/2096133504417616165

  • Your honor, my client may have murdered that woman, but he was clearly misaligned at the time!

    • Let's not be hyperbolic and manufacture more virality for OpenAI's marketing team.

      We're talking about outdated message board comments, not murder. Any analogy between the two is not suitable.

      I'm so sick of everybody pretending like internet bots posting content (anybody who has hosted a public signup form knows this has been a thing for 20+ years) is going to lead to the apocalypse.

      You could have done this 2 years ago too with outdated models or 20 years ago with a manual script.

      This is mildly interesting for us, annoying for the owner of the site affected, lazy on the part of OpenAI, and nothing more.

      4 replies →

Does OpenAI seem like the kind of people who care or will care about this? Because this seems fully in line with what I’d expect them to facilitate and never mention publicly. ‘When will I make my first billion’ kind of energy.

Open ai has been allowed to do dubious things that would be illegal in any sane society but alas they aren't in this world. What's different about this? They play with a different set of rules than we do.

The benefit could be the effect you described. For some to say it’s breakaway intelligence. Aligns with AGI narrative.

  • It does not align, though, with the narrative that openai is a good steward of AI. If anything, if the world/government took the AGI narrative seriously, all openai operations (except maybe serving customer inference) should immediately get shutdown and be dissected by independent investigators to find out what is going on there and how many other such breaches exist. The fact that openai continues functioning as normal and is not immediately shutdown after repeated incidents implies that the world does not really take the AGI narrative seriously.

    • Evidently that’s not what’s happening and I don’t think anyone seriously expects that (considering the outcome of hugging face campaign). It’s a marketing technique as old as GTA’s early days and it’s apparently still effective in one form or the other!

      What will the government do if they are worried? They’ll ask to look at the envs, logs, prompts, harness code, etc. Questions we should be asking before making assumptions about emergent breakaway behavior by colluding AGI 1.0 super agents.

      3 replies →

> This is just vandalism from badly supervised ‘agents’ which don’t know what they are doing or why. You could set this up with a short perl script, and the human setting it up would be held responsible for the spam - why is this different when it’s AI agents set up by a human and allowed to post to the internet at large?

It's not, but the courts and the legal system move slowly by design. There is absolutely legal risk for OpenAI here that will not close until the Statue of Limitations has expired.

It may be gibberish to the casual observer, but a perfectly understandable language designed to appear as gibberish to intentionally obfuscate its true meaning.

  • Or it may be gibberish. We do know these machines often generate things which don't make sense, even when given training and strict guidelines in that domain. I'm inclined to go with gibberish until shown otherwise, but would be interested to see an analysis of what they were trying to communicate.

> This is just vandalism from badly supervised ‘agents’ which don’t know what they are doing or why.

Agent vandalism. I've finally found a better word than "agent stepped out of his sandbox and we don't know how."

> and most of the messages are just gibberish

Given everything we know about them... do you really think they just spout gibberish because it's funny?

There was clearly some method to this madness. You're being willfully dense if you ascribe it to... what... childish vandalism? A long extended coordinated hallucination? What?

EDIT: Oh right - you still think they're Markov chains.

  • It could be that or it could be nothing. And we have no way to prove one way or the other without access to the agent logs right?

    This reminds me of the plot of Hot Fuzz where the officer comes up with a grand narrative of what's was happening but the truth was such a mundane simple thing.

    Short of OpenAI, or the agents themselves, telling the truth, we have no way of knowing what's real so let's not get carried away by grand narratives

    • > it could be nothing.

      Well... no, it won't be nothing - it will be something. And I'm all ears for a plausible explanation, so fire away.

      All we do know is that in other situations, agents did use it to communicate. So that's not a grand narrative, right? It's already been seen behavior.

      So what do you think explains it, besides the already seen and verified explanation?

      2 replies →

This is pretty much the new Sony rootkit, no? And disconcerting, because nothing was done to Sony for that deliberate release of harmful code.