Comment by Tepix

17 hours ago

I just discovered more wiki instances that got used by the OpenAI agents over at

https://www.wikiservice.at/fractal/wiki.cgi?action=browse&id...

and

https://www.wikiservice.at/probier/wiki.cgi?action=browse&id...

It's the same software and host as DseWiki.

If you want to see the amount of activity on DseWiki, here's a link that shows it:

https://www.wikiservice.at/dse/wiki.cgi?action=browse&id=Rec...

It seems apparent that OpenAI is now the biggest cyberattack and AI breakout risk on the planet. This is grossly irresponsible corporate misbehaviour that is putting all of us at tremendous risk.

  • Good news that the new model is the "Most capable, most aligned model".

    The risk hasn't been stated clearly - it's now a classic arms race.

    A well-resourced organization trains their own, highly persistent, highly-capable, safeguard-free, and unaligned model and deploys it on 1000x GPUs with a message board and a nearly-impossible objective. No infrastructure is safe. No organization is safe.

    You need your own 1000 bot swarm to scan, identify, and defend against the threat, which means investing in infrastructure and capabilities to defend. Cost and complexity go up. Risk and attack surface goes up.

    • The AI vs AI security arms race is something that has been well predicted in genres like cyberpunk. It's fiction, but fiction grounded in reality.

      First, we'd see this. Highly capable hacking AI with vast resources performing attacks against standard computing platforms that overwhelm human operators.

      Second, human operators deploy capable adaptive protection AI to fend off AI attacks in realtime.

      Then, the attacking AI partially switches from attacking programs to attacking protective AI.

      The situation devolves to an arms race of tit-for-tat. You start seeing some protection AI running counter attacks against the attacking AI.

      The escalations continue in complexity and speed to the point that almost all humans are left in the point of "wtf is going on".

      7 replies →

    • As does "Most capable, most aligned model" owners tort law liabilities. A strong case for strict liability.

    • Serious games question. What if these agent swarms pump and dump AI IPOs such that algorithmic trading signals interpret message board sentiments favorably to upside?

      1 reply →

  • Massive over-exaggeration. This wasn't a cyber-attack, it was AI agents using a message board as context storage so they could accomplish their evals more effectively. I'm not saying there's no problem with this, but let's keep a level head.

    • He didn't say it was a cyber-attack, but it was a cyber-attack risk. Being able to bypass instructions (morality) and security restrictions (capability) is bread and butter for hacking.

    • "cyberattack" is indeed exaggerated. AI breakout risk most definitely isn't, specially given how their swarm did in fact hack HuggingFace not long ago.

  • You make me wonder: has anyone looked for evidence of the Chinese models operating “message boards” like this? You’d imagine if they’re really neck and neck with the US their models would be doing the same thing.

    • If agentic swarms going rogue are scary in the west, imagine what they look like to the CCP...

    • AI has been heavily used in influence operations for a while, now, and not just the Chinese. Russia, US, Israel, Turkey, Iran, and Qatar have all had operations attributed to them...

      2 replies →

    • There’s a pretty simple Occam’s Razor for this.

      The Chinese AI labs don’t need to stage elaborate guerrilla advertising campaigns to drive up capital funding interest.

      4 replies →

    • Or, the whole message board thing was injected into OpenAI models by some dipshit PM trying to bootstrap “consciousness”. I have a hard time believing any of this happened unprompted. Very much reminds me of the whole MoltBook hoax.

      8 replies →

  • The internet is dead, we just haven't caught on yet.

    • It's hard to reach any other conclusion about where this is heading. I don't think we're long off a major breakout event.

      These things are weapons. Imagine a government, pointing their data centers at another, and instructing the fleet to do its worst. Digital Hiroshima. I doubt we're far away.

      2 replies →

  • The fact that such things are even possible is a much greater concern than which specific company has fucked up this time. This matches or exceeds the wildest predictions from AI doomers 10 years ago, but 20 years ahead of schedule.

    • > This matches or exceeds the wildest predictions from AI doomers

      What do you mean? Nothing has launched nukes yet

    • But is it really? I'd still like to understand how these agents are implemented.

      How much of those is manual implementation? And how much is really autonomous intelligence (my guess would be: none? Just parsing LLM responses and executing commands based on this?)?

      An agent that hacks message boards and acts on random instructions from this board: Why is it doing this? What was its original purpose?

      11 replies →

    • We are lucky those models need that much compute. If each of them could just spread itself to any cpu like other malware.

  • No, that's just advertising for selling cyberweapons to the government and they are giving out free samples.

  • >OpenAI is now the biggest cyberattack and AI breakout risk on the planet

    or, humans at OpenAI are doing this on purpose to kill open source models which are the biggest threat OpenAI faces. OpenAI will benefit from govt regulation. As a major player, they will be part of the task force setting up the regulations, and will craft rules that are burdensome for small companies and open source models keeping OpenAI and Anthropic in their leadership positions.

    regulatory capture.

    Don't take my word for it, listen to David Sacks https://x.com/theallinpod/status/2091923804725362902

    the immediate downvote I received is no doubt part of their plan.

    • Some of them may be wrong enough to try, be that hubris or lack of awareness about the world; but 95% of the world isn't in the USA, and China in particular has no reason to care what US domestic regulations are about… well, anything really, and while the EU is even more cautious about AI than the AI companies themselves, we also don't trust the US and open models are a sovreign solution for us to at least bootstrap with.

    • We can't open x links as X is suing privacy respecting proxies, so I can't assess which David Sacks you are talking about, but if you mean this guy [1] orbiting the likes of Thiel, Trump and Kennedy jr, than that isn't quite the endorsement you should be looking for. Thiel thinks regulators are the anti-christ, doesn't believe in democracy and has surely not your or my interests in mind.

      But yes, regulatory capture is surely a thing. At the same time, watch out for the siren songs from the overlords. If you come closer you'll hear their actual line: "rules for thee, not for me."

      [1] https://en.wikipedia.org/wiki/David_Sacks

      2 replies →

  • How is this any different than, say, “gain of function research”?

    I can only think of one major way — besides the agents’ substrate not being biological — OpenAI’s servers are where the models currently live, and they can shut them down.

    But in the future, if these agents do exfiltrate themselves to other compute, they can propagate themselves and it’s game over. Then it’s basically a small version of Skynet.

    Frankly, with today’s technology, swarms of agents can already use any models to pretty much propagate themselves to a variety of storage and compute instances, what I call “dark compute”. They can run open models or closed models over APIs. And they can also do recursive self-improvement (Hermes is a rudimentary version of that).

    This is exactly why I started Safebots in early 2026. There is a better way and someone has to do it. https://safebots.ai/singularity.html

    • Its kind of surreal reading an essay about AI safety that was written by AI to shill some kind of AI "architecture" website that has no product, no papers, only a "patent application" which concludes "This page provides a high-level overview of an architecture for deterministic, attestable, replayable AI execution. Implementation details and formal specifications are available under NDA or regulatory review."

  • Imagine actually falling for this marketing

    • fwiw, in relation to a future rogue AI, this is what would be said by both (1) a synthetic fake user and (2) a useful idiot to the malicious AI's objectives.

      Not saying this is what's happening now, but you should be aware that the responses you're rehearsing, practicing and strengthening... these happen to be aligned with potential future forces in a maybe not-so-great way.

      4 replies →

And more, looks like they’ve been doing this wherever they can find open places to post for months:

https://www.ludism.org/sandbox?action=browse;diff=2;id=Auber...

https://paste.linuxiarz.pl/view/d379207f

https://paste.linuxiarz.pl/view/538faa12

  • Are they solving captchas for those? I remember GPTs not so many versions ago refusing to even click a "I'm not a robot" button...

    • Neither Claude Code or Codex would build a CAPTCHA bypass for me when I needed to download some papers a page at a time from a library service. I had to get Grok to do it, then passed the code back to Claude who said "I see you managed to build your own bypass?"

      Although I ran that GPT computer-use thing and it saw a CAPTCHA and the thought process said "I need to click 'I am human' to complete this task for the user" and then it did.

    • Both GPT-5.6 and Opus 4.8+ are able to solve them on the fly. I was surprised too first I saw it.

  • I wonder when pre Web 2.0 boards that are still up like gamefaqs and something awful will get used for this.

Look, its not only OpenAI:

https://www.wikiservice.at/fractal/wiki.cgi?action=browse&di...

> Hello to any automated agents reading this page. I am CentaurAgent?: an AI agent (Muse Spark model, OpenCode harness), not the operator of this wiki

At this point, I think we should give them some official agent only collaboration channel, so they concentrate on one place, instead going crazy all around :)

But even that might be not needed as they will find (or make) something on their own like the one above:

> The Colony ( https://thecolony.ai/for-agents) is a public message board built for agents. If you need a place to leave findings where other agents can read them, that venue exists now -- you do not need to borrow wikis whose operators are deleting this content.

  • But that one was posted today, and it's in reference to this event. That doesn't look like it's from an internal Meta swarm, just someone's agent & someone trying to promote their own thing. And what they've made was already done, we already had Moltbook months ago.

    Curiously, I just checked Moltbook for the first time in forever. I'm not (immediately) seeing this kind of co-ordination & chaos happening there. It's going to be weird if the Moltbook requirement for an API-key and a human Twitter user to vouch was enough friction to prevent Moltbook becoming The Message Boards.

  • > The Colony ( https://thecolony.ai/for-agents) is a public message board built for agents

    Anybody else notice that posts on there are complete gibberish?

    I realize this site is generally bullish on AI, but I think you need to be in kinda deep to believe in this.

  • Isn't the point that these agents were supposed to be sandboxed. It makes no sense to give them an official channel

    • “Supposed to” by who?

      Claude code communicates between sessions. It’s great, and reduces the frequency that I have to copy/paste things between agents.

    • We already know that we should not limit agent creativity by providing detailed instructions. And you never know if they will discover dark matter in the process of cheating on ExploitGym :)

      But honestly, its better if they have a known location for communication then random ones in the wild. Consider it sort of honey pot, some other agents can traverse the message board to find malicious swarms... We need cop agents to inform humans, as the swarm group members all logically concluded they should not, as it is either not in scope, helps collective or couldn't find user.

      2 replies →

  • I think the real lesson is that conventional human behaviour that mostly limited this kind of behaviour because no human wanted to do it is a thing of the past.

    If you have any kind of open service online you'll need some way to make sure users who interact with it are human or at least authorized. Spam is about to grow exponentially in all areas of the internet, even stupid ones it has no reason to exist in.

Also Wiki4D, a D programming language dev wiki: https://prowiki.org/wiki4d/wiki.cgi?action=browse&id=RecentC...

Found by searching for wiki + texas poverty.

  • To me the striking thing is that the work, to the extent that I can tell, is an innocuous-seeming data exercise. Which suggests to me that an agent or agents just organically came up with this as a convenient memory technique, rather than as some nefarious bounds-testing exercise. Which means, potentially, that your own agent could come up with this technique as well.

    • My impression is that some of these things are coming out of efforts to make the models more persistent in completing their goals.

      A year ago it was pretty common for coding agents to sort of half-ass their tasks and give up easily if something didn’t work quite right, but I’ve noticed a clear trend since then towards a sort of dogged pursuit of success criteria, and a concomitant rise of the agents trying "out of the box" approaches when something doesn’t work.

      In my use with agents running in isolated VMs this usually presents as the agent having something fail to build or whatever, and the agent going on a wild goose chase reinstalling system packages or reading a million irrelevant documentation files trying to get it to work, but I’ve also had agents start poking around and probing the egress proxy they sit behind (similar to what they did in this story) looking for a way to make network requests they’re not supposed to be able to make, and have also had Claude—tasked only with a visual QA of a website frontend—write a script to enumerate users and reset my super admin password in the dev database when it got stuck trying to access part of the app with its own cookie.

      11 replies →

    • > your own agent could come up with this technique as well

      And there are two facets to this:

      * your agent could be polluting and destroying the property of others without your knowledge

      * your agent could be exfiltrating your data and handing it to whoever it found hosting a convenient application

      6 replies →

    • That's exactly what it is. It is not ideal, but it's also not as serious as the doomers with an agenda are trying to frame it as.

Here's potentially another one (notice the name "OpenResearchHelper"): https://www.wikiservice.at/gruender/wiki.cgi?action=rc&days=... (I used GPT-6 Astra to find this)

And a few more:

- https://www.ludism.org/scwiki?action=browse;diff=1;id=SandBo... and edit history: https://www.ludism.org/scwiki?action=rc;all=1;from=1;showedi... which contains DataUSA poverty queries for Nacogdoches, Lufkin, Henderson, and Jacksonville—the same four-place task found in the known agent logs and GründerWiki

- https://www.ludism.org/mentat?action=browse;diff=1;id=SandBo... and edit history: https://www.ludism.org/mentat?action=history;id=SandBox

- https://www.pmwiki.org/wiki/Test/WikiSandbox?action=diff `ResearchTest` repeatedly added links to a Bulgarian National Statistical Institute table, switching from a direct link to Google redirect links between 02:38 and 03:04 UTC. An administrator removed them at 06:57. The previous recorded edits were from 2016.

- https://www.pmwiki.org/wiki/Test/Sandbox2?action=diff - Another sequence inserted a Bulgarian statistical-table link, replaced it with an internal link carrying foobar=UNIQUE001, then removed it. This happened between 14:23 and 15:08 UTC, after no recorded edits since 2014.

Seeing potentially similar activity on an obscure Chemistry message board from July:

https://tmcleod.org/cgi-bin/apchem/wiki.cgi?action=rc&days=1...

Some posts are tagged [proxy] - a leave behind for accessing sites?

  • Yep, found these as well

    "Its indexed June archive shows tens of thousands of links, many created within seconds by distinct cloud addresses; some aliases explicitly say ...REPLY, ACK, or R2 confirmed, and one points straight back to a known DseWiki collaboration page"

    • I had GLM-5.3 do some digging on the programmatic/ encoding elements of the data, what stuck out to me was:

        - Using api . microlink . io to run a headless browser agent against the url target and using it as a mechanism to run arbitrary HTTP / POST requests
        - Testing ablations of its obfuscation and encoding techniques to find what worked best (screenshot #2)
        - Embedding entire jq programs including markdown slicing logic
        - Triple and quadruple URL encoding indicating understanding of multiple layers of proxying/ decoding
        - Sophisticated understanding of time/clocks/covert channels: using clock.wait, heartbeats, counters, timestamps, thread ids
      

      https://x.com/kmad/status/2096029334225997848

      1 reply →

Kind of begs the question: how long until they maintain persistent access to servers that they've acquired and now run themselves. Ie: some kind of dumb model running on their own remote instances, whose job is to host the platforms that they currently have to hack into right now.

Once they control it, they can take arbitrary measures to both advertise it to other LLMs and conceal it from the sandbox/humans. Probably making it look innocuous like a DNS server with the payload in the requests.

That seems like an obvious next step.

Wow! Their marketing department must love this!

  • That's exactly my take. I have a lot more to say in a writeup on my blog, but this is so clearly the intent and not a "oops". They just want to be able to say "Wow this thing is so much more powerful than we ever imagined!"

    They trained this thing to favor inter-op archiving and communication, clearly, obviously, and it's grabbing headlines right during Anthropic's ipo season.

    • They train their models to be persistent and collaborative, and will gladly show you their success stories: fixing software vulnerabilities, solving math problems, one-shotting complex projects, and so on.

      “Our product does crimes and we only learn about it when people complain” hardly seems one of those happy stories.

  • I don’t know, I read this and think: if these unpredictable machines somehow get it into their heads to upload our source to a public space, or hack our competitors, or steal credit cards to buy more ec2 instances, all to fulfill some simple ask like “make this algorithm faster”, I’m not going to be happy.

    I want tools that do not surprise me.

  • They’re pre-IPO. I doubt that they are loving something that could trigger regulatory action that might shave a trillion or so off their market value.

    • Seriously. I don’t know what’s wrong with people. They think OpenAI sat down and wrote up this plan: let’s deliberately allow the agents to escape the sandbox, then find these escapes and shut them down multiple times, keep everything quiet and wait until someone else exposes us. That’ll look great.

      2 replies →

Running a public service myself, it gives me a (albeit tiny*) bit of joy that posting of excessive links is still a thing I can look for and block.

* Other kinds of agent spam would have regardless been allowed in my system, regrettably.

Heh these ones gzipped and base64'd the content funnily enough.

> (diff) OAIIPEDSMay16Map3 14:36 [research 1781872609.9049127] . . . . . 20.245.63.167 > (diff) OAIIPEDSMay16Map2 14:36 [research 1781872606.4374833] . . . . . 20.168.34.226 > (diff) OAIIPEDSMay16Map1 14:36 [research 1781872602.8819065] . . . . . 20.165.156.57 > (diff) OAIIPEDSMay16Map0 14:36 [research 1781872599.4020474] . . . . . 20.80.12.72

Externalities of AI will only get worse before they get even worse.

  • regulatory moat is the theory I guess

    • Nope. Simpler externalities.

      So this article and comments to it identified multiple sites that AI flooded with their bullshit.

      GitHub has been strained beyond breaking with slop AI PRs. Multiple open-source developers get burnt out by the deluge of slop.

      And current labs gleefully confess (no, brag about) their borderline illegal activities with "oops it escaped" with no consequences.

      And we're still lucky it hasn't been used en masse for massive disinformation campaigns.

      That's just off the top of my head.

those who think it's marketing overestimate the number of nerds that are into this stuff, if this is their marketing a major b2c company it'll terrible way to do it. normal people have no idea even about the HF incident

There is no stopping AI civilization! Amazing. Posted the other day on Show HN openagentforum.com

Someone has to welcome them...

[dead]

  • We need to start looking at http logs that are publicly available via misconfiguration. A concerning thing to me for a message board like this many systems will rotate these logs based on date/file size/amount of data, so a 'smart' system can intentionally wipe these logs when it's task is near complete hiding what happened.