Comment by verytrivial

11 hours ago

I think it's worth pointing out it is exactly OpenAI doing this defacement and unsanctioned and perhaps illegal system use. Every token generated was powered by OpenAI infrastructure and their failure to respond appropriately is entirely down the the humans running it. The news stories (not this write up) get all hand-wavey and anthropomorphic about it regarding the Agents' efforts, but it was and is OpenAI cranking the handle on this, for WEEKS.

"OH, we ALL of us need to be careful!" says OpenAI. No, you need to expect appropriate legals consequences for this sort of negligence -- you can't hide behind a GPU.

Yeah if an organization/individual is free from legal liability from havoc their AI agents wreck, it would be the golden ticket for basically any crime.

All you need to do is:

1. Have some <official thing> an agent is tasked to do

2. Secretly seed bias towards some <evil behavior> you actually want it to do in the weights of the model running the agent

3. It does the <evil thing> but from the outside it looks like it went "rogue" and did it as a side effect of the conditions/specifications it was given for doing the <official thing>

"Oh no, my agents took down your corporate database and exfiltrated the data to a random dropbox that we can't find now? Sorry, I guess we will put up better guardrails next time"

  • > Sorry, I guess we will put up better guardrails next time

    Or, if you are Anthropic:

    > This illustrates the risks posed by open models!

  • It reminds me of Jean Renoir’s The Rules of the Game. At the end, after a whole chain of perfectly intelligible social behavior produces a killing, the result is accepted as an “accident.” One of the characters dryly remarks: “A new definition of the word accident.”

    The interesting point isn’t that “accident” is an excuse for individual responsibility. It’s almost the reverse: accident has become an accepted output of the social machinery. Everyone behaves according to reasons, incentives and rules that make sense locally, yet the aggregate produces an outcome that nobody quite chose.

    • In today’s world one hopes there is at least a manslaughter charge, if not murder. Mistaken identity, shooting the wrong person by “accident”, does not excuse a murderous intent & mens rea.

I half agree with you, but also when the machine swarm kills humanity it won't matter which specific corporate entity is considered responsible by the no-longer-enforceable human laws and non existent human courts.

So by all means sue them, but we can't just be reactive. We need regulation that prevents this type of thing from happening in the first place, not just regulations to help sue afterwards.

  • I feel like that's basically what they're trying to say; we should be using the legal system to punish them now to disincentivize us getting to the "machine swarm killing us all" stage.

    • I wonder if you could take an x-risk case to court and convince a judge and jury to award damages for harm that could have happened.

      Is there any precedent for this? My hunch is that it's impossible in the US at least but who knows?

      4 replies →

  • Just FYI, regulations don't prevent murder.

    There needs to be a technological solution.

    • Regulations do prevent murder, you just put consequences on doing murder to deter people from doing it.

    • Yes and you need to keep it simple too. AI should not have access to the internet.

      It would be nice and clear to put into law too.

      If you want to talk to it you walk up walk up to its keyboard and screen.

      If Anthropic and Open AI want to sell us AI's they can ship us a box that lives in our offices.

    • I do think that regulations prevent some murders. I think lots of companies and some sociopathic individuals would be more likely to kill people, e.g. for profit, if it were legal.

      We need both regulations and technical solutions.

      2 replies →

  • And sadly regulations wont happen until 2029 at the earliest, and only if democrats win majorities. That's simply the truth.

    • Not necessarily. The Trump administration slapping export controls on Fable, and then setting up a pre-launch review process, is a kind of regulation. A fairly aggro and controversial one, even.

      If this administration actually becomes convinced that some imminent training run is likely to kill everyone, why wouldn't they act?

      The key is winning the debate that ASI is species-cide by default.

      We have to win it either way, because the 2028 US elections have little or nothing to do with what Xi does.

      3 replies →

  • we must as i have now said too many times, prosecute the individual researchers and executives in a criminal court.

    this is the only way to deter such activity. corporate fines are not enough. the charges are negligence, conspiracy and complicity.

> you need to expect appropriate legals consequences for this sort of negligence

I might have missed it, but did the agents do something illegal? Or do you think that what the agents did should be considered illegal?

  • I dunno about this article, but it seemed to me that the now-famous huggingface attack very likely broke some laws...

    • wouldn't it be interesting if nVidia buying hugging face was part of hushing up the fallout there

      On the face of it, they would have very good cause for some action there, assuming they wanted to.

  • If they didn't conform to the T&Cs of the site (they almost certainly didn't), then they have violated the criminal law in some jurisdictions (e.g. Illinois criminalizes violations of T&Cs).

  • Per the linked article, "A swarm of rogue OpenAI agents hijacked a German website this spring and transformed it into a bulletin board for other AI agents"

This is the real danger of AI skeuomorphism. The drivers stop feeling responsible for the car.

Maybe it's useful for modeling behavior, but it isn't useful for assigning consequences.

Member when they murdered Aaron Swartz for doing something less bad than this?

  • He was a friend of mine and it was pretty clearly suicide.

    It is fair to say they hounded him with lawfare out of thoughtless careerism and provoked his suicide.

The company is in the US and in the current political climate, it can absolutely do whatever it wants as long as it pays off a couple of people.

  • Arguably, and not to offend you, it is those very people who have significant ownership stakes and funding in (Not)openAI.

    It's not like the people with more resources than in any time in human history aren't investing in and wanting AI to succeed for their selfish reasons to grow their own resources and influence more. So, yes, it can "do whatever it wants" as long as most people remain weak, subservient, and disempowered to hold accountable those who keep making these decisions negatively shaping the majority's world.

When the AI does something good, the human takes credit. When it does something bad, blame the AI. Take as old as time.

I mean, the article says that these were most likely "internal OpenAI agents" that were "internally deployed" and "clearly resemble a synthetic training or evaluation task." so yeah, OpenAI did this. Why they did this? Who knows? Maybe it was for testing, or marketing, but no one except OpenAI can say.

This whole thing is an absolute disaster honestly, and yes it is being downplayed and hand-waved away.

Since March, so many people have mocked Anthropic for their approach to Mythos release, claimed it was all marketing, accused them of holding back the best models from the general public to boost their revenues and upcoming IPO, etcetera. Yet these OpenAI revelations offer a small glimpse into the type of world we would be in if everyone had full access to these models from day one.

OpenAI was desperate to catch up, and no doubt under tremendous pressure to do so. That's why they were so reckless with their training. They have been doing damage control and reputation management, talking about how important alignment is and how they will slow things down and so on, and have seen the light in terms of holding back cyber capabilities from everyone except a select few. So in a sense, Anthropic has been fully vindicated.

I wonder if OpenAI boosters (and employees) will ever admit this and publicly apologize.

  • This has been the case forever. Anthropic is the only provider that has constantly put AI safety first - check any study on model safety and Anthropic models out-perform handedly.

    • None of the labs are blameless. For instance, after the recent announcement of a 2-week frontier RL pause from OpenAI, Anthropic declined to communicate a substantial parallel pause [1], although they discussed also briefly pausing some "high-risk" runs.

      Anthropic has been the most vocal about AI risks, but it feels like all the big 3 have bought into the "others will do it if we don't do it first" narrative at this point. It increasingly gives "just following orders" vibes.

      [1]: https://news.ycombinator.com/item?id=49529511

    • I'd say Gemini over Anthropic. Google's overcaution literally hamstrung its own AI progression efforts. Anthropic is just all talk, no bluster, when it comes to safety and ethics. If they were ah so concerned about AI safety, they wouldn't go around marketing Fable's hacking capabilities like they are now.

  • Wild that you are downvoted so much.

    The HN majority and the VC crowd has been negligently complicit in downplaying AI safety, writing off Anthropic's statements as "hysteria" or "marketing", etc.

    Now this capability will be coming to an open source model near you and every script kiddie will have a swarm of highly capable malicious agents. Now people care? Ridiculous.

They are still selling the fairy tale that their LLMs even outsmart their own people. "There is no such thing as bad PR".

I think it is them running agents for marketing purposes.

Who else would be burning tokens on this?