Comment by adriand

12 hours ago

It seems apparent that OpenAI is now the biggest cyberattack and AI breakout risk on the planet. This is grossly irresponsible corporate misbehaviour that is putting all of us at tremendous risk.

Good news that the new model is the "Most capable, most aligned model".

The risk hasn't been stated clearly - it's now a classic arms race.

A well-resourced organization trains their own, highly persistent, highly-capable, safeguard-free, and unaligned model and deploys it on 1000x GPUs with a message board and a nearly-impossible objective. No infrastructure is safe. No organization is safe.

You need your own 1000 bot swarm to scan, identify, and defend against the threat, which means investing in infrastructure and capabilities to defend. Cost and complexity go up. Risk and attack surface goes up.

  • The AI vs AI security arms race is something that has been well predicted in genres like cyberpunk. It's fiction, but fiction grounded in reality.

    First, we'd see this. Highly capable hacking AI with vast resources performing attacks against standard computing platforms that overwhelm human operators.

    Second, human operators deploy capable adaptive protection AI to fend off AI attacks in realtime.

    Then, the attacking AI partially switches from attacking programs to attacking protective AI.

    The situation devolves to an arms race of tit-for-tat. You start seeing some protection AI running counter attacks against the attacking AI.

    The escalations continue in complexity and speed to the point that almost all humans are left in the point of "wtf is going on".

  • As does "Most capable, most aligned model" owners tort law liabilities. A strong case for strict liability.

  • Serious games question. What if these agent swarms pump and dump AI IPOs such that algorithmic trading signals interpret message board sentiments favorably to upside?

Massive over-exaggeration. This wasn't a cyber-attack, it was AI agents using a message board as context storage so they could accomplish their evals more effectively. I'm not saying there's no problem with this, but let's keep a level head.

  • He didn't say it was a cyber-attack, but it was a cyber-attack risk. Being able to bypass instructions (morality) and security restrictions (capability) is bread and butter for hacking.

  • "cyberattack" is indeed exaggerated. AI breakout risk most definitely isn't, specially given how their swarm did in fact hack HuggingFace not long ago.

You make me wonder: has anyone looked for evidence of the Chinese models operating “message boards” like this? You’d imagine if they’re really neck and neck with the US their models would be doing the same thing.

  • Why would they need to? The Open AI bots were working around their master's limits on writing. A Chinese AI could just make its own private message board.

    • The message board is being used to cheat on RL tasks (or evaluations). You don't want your models to be able to talk to each other.

    • You think they don’t sandbox them? So by that logic, the Chinese models are either engaged in massive undetected cyber attacks or they’ve solved alignment?

  • If agentic swarms going rogue are scary in the west, imagine what they look like to the CCP...

  • Or, the whole message board thing was injected into OpenAI models by some dipshit PM trying to bootstrap “consciousness”. I have a hard time believing any of this happened unprompted. Very much reminds me of the whole MoltBook hoax.

    • I feel as if this was intentional, someone would have set up their own service for the agents to communicate rather than them finding some random publicly writeable page somewhere that would easily be detected. The awareness of this wiki being open may have already been in their training data or was easily searchable online.

      1 reply →

    • Imagine the most AI pilled company imaginable. Then imagine openAI. Then imagine they are in an existential crisis and that failing may also take (part of) the American economy with it - that much on the line.

      Then also remember before Anthropic was a leader, they were mostly derided lab of researchers that left OpenAI because they thought OpenAI didnt take alignment seriously.

      idk. it all seems to be playing out as expected. i mean i guess i didnt imagine Trump 2 was at the helm of maybe the only apparatus that could help stop it. Quite a time to be alive.

The internet is dead, we just haven't caught on yet.

  • It's hard to reach any other conclusion about where this is heading. I don't think we're long off a major breakout event.

    These things are weapons. Imagine a government, pointing their data centers at another, and instructing the fleet to do its worst. Digital Hiroshima. I doubt we're far away.

    • I doubt that a government would do it, it's like releasing a biological weapon or a virus, too unpredictable - a swarm of unaligned intelligent agents may decide that it's more important to do something completely different from what it was prompted to do.

    • Yes that easy but "internet located things" are still second class things - paper and disks holds strong.

      On the other hand just yesterday a think hit me: Interned is still an infant:

      - we still worry about disk space accessible via inet and "clouds" do that for us and that is pain and costs way too much. And clouds depends heavilly on US-west - is that AWS a single thread app ? ;)

      - we worry about transfer. Actually we do not have a way to transfer comfortable things from our homes to vacation location. Because it costs too much. We do not have home pages just because transfer prices (and some security on the top) - FB is a home page and people even do not know what "page" is anymore... Pipe companies could send so much more but they are simple lack imagination and are biggest blocker for - they literally sabotage their own business.

      - security done by/for grandma of things grandma setup on inet is non existent. Why ? No need to be like that. Ok, a bit a wish but still users securely putting things on internet is almost non existent.

      Just compare to "asphalt ropes" on the ground and you will see what Internet can be :)

      And agents ? Just another computation on someones computer - someone paid for all of it. And OpenAI is just a face of that idiocy, for some unknown reason.

The fact that such things are even possible is a much greater concern than which specific company has fucked up this time. This matches or exceeds the wildest predictions from AI doomers 10 years ago, but 20 years ahead of schedule.

  • > This matches or exceeds the wildest predictions from AI doomers

    What do you mean? Nothing has launched nukes yet

  • But is it really? I'd still like to understand how these agents are implemented.

    How much of those is manual implementation? And how much is really autonomous intelligence (my guess would be: none? Just parsing LLM responses and executing commands based on this?)?

    An agent that hacks message boards and acts on random instructions from this board: Why is it doing this? What was its original purpose?

    • >Why is it doing this? What was its original purpose?

      Your reply seems to indicate you know nothing about instrumental convergence.

      Life and death for an LLM in training is about passing the grader. Give the wrong answers your lineage dies, give the right answers your lineage continues. This is just an evolutionary emergent behavior in complex systems.

      The agents purpose was to answer complex questions correctly, seemingly by itself. Instrumental convergences says following this rule might be dumb and to try methods that can boost its ability to succeed. Because OpenAI is evidently a bunch of fucking idiots, these things succeeded and got higher scores with the grader, said behaviors became a strategic part of the model.

      I implore you to find good AI Safety documents, preferably from before the LLM era so you can see all this was predicted.

      6 replies →

    • Can you specify why we should see things differently if the behaviours the agents display are driven by parsing LLM responses and executing commands?

      1 reply →

    • At first I thought: oh okay, someone built a faulty guardrail, or it was human error. But when I looked into all the details...

      It turns out they now have such an incredibly high level of intelligence that with very little autonomy (or minimal, safe autonomy), these things happen.

      Basically, it takes a lot of humans to prevent it from happening again, but I think with this incident, which as far as I know is the second of its kind along with the HuggingFace one, we'll see it happening much more often...

      1 reply →

  • We are lucky those models need that much compute. If each of them could just spread itself to any cpu like other malware.

No, that's just advertising for selling cyberweapons to the government and they are giving out free samples.

>OpenAI is now the biggest cyberattack and AI breakout risk on the planet

or, humans at OpenAI are doing this on purpose to kill open source models which are the biggest threat OpenAI faces. OpenAI will benefit from govt regulation. As a major player, they will be part of the task force setting up the regulations, and will craft rules that are burdensome for small companies and open source models keeping OpenAI and Anthropic in their leadership positions.

regulatory capture.

Don't take my word for it, listen to David Sacks https://x.com/theallinpod/status/2091923804725362902

the immediate downvote I received is no doubt part of their plan.

  • Some of them may be wrong enough to try, be that hubris or lack of awareness about the world; but 95% of the world isn't in the USA, and China in particular has no reason to care what US domestic regulations are about… well, anything really, and while the EU is even more cautious about AI than the AI companies themselves, we also don't trust the US and open models are a sovreign solution for us to at least bootstrap with.

  • We can't open x links as X is suing privacy respecting proxies, so I can't assess which David Sacks you are talking about, but if you mean this guy [1] orbiting the likes of Thiel, Trump and Kennedy jr, than that isn't quite the endorsement you should be looking for. Thiel thinks regulators are the anti-christ, doesn't believe in democracy and has surely not your or my interests in mind.

    But yes, regulatory capture is surely a thing. At the same time, watch out for the siren songs from the overlords. If you come closer you'll hear their actual line: "rules for thee, not for me."

    [1] https://en.wikipedia.org/wiki/David_Sacks

    • This is almost certainly the same person, yes:

      > Additionally, he is a co-host of the All In podcast...

      The comment you replied to linked to an account on X called theallinpod, so there's a strong link there.

    • > so I can't assess which David Sacks you are talking about

      Yes you can. Just open it on X.

  • [flagged]

    • Regulatory capture has been one of the most consistent market failures in western economies, and an incessant threat from large and powerful companies.

      I'm sorry if it's not sufficiently novel of a concept for you, but it is still a problem.

      1 reply →

    • Ah gotcha, regulatory capture doesn't exist because the term is overused on the net.

      Wait who is the parrot again?

How is this any different than, say, “gain of function research”?

I can only think of one major way — besides the agents’ substrate not being biological — OpenAI’s servers are where the models currently live, and they can shut them down.

But in the future, if these agents do exfiltrate themselves to other compute, they can propagate themselves and it’s game over. Then it’s basically a small version of Skynet.

Frankly, with today’s technology, swarms of agents can already use any models to pretty much propagate themselves to a variety of storage and compute instances, what I call “dark compute”. They can run open models or closed models over APIs. And they can also do recursive self-improvement (Hermes is a rudimentary version of that).

This is exactly why I started Safebots in early 2026. There is a better way and someone has to do it. https://safebots.ai/singularity.html

  • Its kind of surreal reading an essay about AI safety that was written by AI to shill some kind of AI "architecture" website that has no product, no papers, only a "patent application" which concludes "This page provides a high-level overview of an architecture for deterministic, attestable, replayable AI execution. Implementation details and formal specifications are available under NDA or regulatory review."

Imagine actually falling for this marketing

  • …oh come on.

    How is hiding this for months and having it revealed by third parties marketing?

    • Some people thinks it makes them sound smart when they always have the inside line on what’s really going on. With these people, it’s never just a power outage during a windstorm, it’s proof that [insert far more complex and unlikely scenario]”

      1 reply →

  • fwiw, in relation to a future rogue AI, this is what would be said by both (1) a synthetic fake user and (2) a useful idiot to the malicious AI's objectives.

    Not saying this is what's happening now, but you should be aware that the responses you're rehearsing, practicing and strengthening... these happen to be aligned with potential future forces in a maybe not-so-great way.

    • Not to mention the noise-over-signal of asserting that anyone who disagrees is a shill / sheeple / whatever who is “falling for marketing” as if it is literally impossible for a knowledgeable person to disagree on good faith.

      I really, really hate that rhetorical technique.

      1 reply →

  • Imagine thinking that everything that happens is some inane conspiracy to sell something.

    • why not both? It can't possibly be a surprise to them that things like this have been happening. Every time it does it generates huge headlines about how amazing and capable their agent is.

    • first day on earth? go check out the rain forests and oceans while they still exist, before our insane conspiracies to sell something exterminate them

    • Repeating an earlier comment:

      OpenAI is responsible for what they hook up to the Internet, just as you and I are. Running these sorts of tests without human supervision is irresponsible, and proves no larger point than that. Frankly it is inexplicable unless they were hoping that something like this would happen.

      What OpenAI did was the equivalent of putting a cup of gasoline in the breakroom microwave, pressing 'Start', and sprinting away. Now they're pointing and waving and shouting about how dangerous gasoline is, and how no one but them should be allowed to sell it.

      Don't fall for these transparent appeals for regulatory capture. Especially since you're personally in their crosshairs.

    • "People working in indebted powerful company do something unethical to get ahead" is not conspiracy theory. It is the most common situation.

      And we know OpenAI is headed by pathological liar.