← Back to context

Comment by damowangcy

17 hours ago

Imagine having a virus escape a sandbox, why are we worried about the virus but not the incompetency of those who are responsible for setting up the sandbox?

If I post something on the Internet today claiming that I asked my agent to do X but it went rogue and did Y, all I will be getting in return is a jar full of "skill issue".

Should we worried about people using LLMs for attacks? Yes, but not in the premise of LLMs going rogue but someone with the intention of abusing it to cause harm. And this is not something we as individual or even company can deal with, responsibility should be held by those who use it, in a legal way.

I am baffled by the fact that up until now, no one is held responsible for so many incidents reported publicly or privately. At this point, it's free marketing, if I am CEO of any AI company, I will run swarm of agents hacking all NGOs and stating that I am just looking for some random piece of data that happened to be hidden in their servers, at least that's what my LLMs think, not me. Then I will start preaching everyone how dangerous this piece of technology is and start giving out free tokens for these NGOs so they can start defending themselves and we should slow the f down.

> Imagine having a virus escape a sandbox, why are we worried about the virus but not the incompetency of those who are responsible for setting up the sandbox?

> Should we worried about people using LLMs for attacks? Yes, but not in the premise of LLMs going rogue but someone with the intention of abusing it to cause harm.

[Why-not-both?-meme]. To use your example, when you discover prions (a class of pathogen that is much more robust to standard disinfection methods than viruses) you should both be worried about your concrete outbreak of BSE (UK in the 80s and 90s) as well as the wider implication (e.g. do we need to change the sterilization methods for our surgical instruments?).

Seriously, I find the way these discussions are done to be super frustrating, because often people implicitly form tribes that oppose everything the other tribe says. When someone believes AI companies push greatly exaggerated stories of dangerous rogue AI to force out competition via regulation they often implicitly conclude that their argument is fundamentally wrong, whereas in reality the lies that work best are those that distort the truth.

Companies should be punished harshly for the deeds of their AI agents AND we should not allow them to force out competition AND we need take the threat of autonomous AI agents as a new class of danger serious AND we need to worry about the socioeconomic implications of AI companies privatizing new means of production.

Yes, there is competition of these ideas in the attention of the general public, but the methods we can use to solve these problems don't compete with each other. AI slowdown for example helps with all the other topics.

  • Couldn't agree more. We should be worried about both things.

    But I share the original posters bafflement that the mainstream conversation seems to accept that framing that the agents were independent intelligences rather than computer programs that the organization that created them is responsible for.

    • > that framing that the agents were independent intelligences rather than computer programs that the organization that created them is responsible for.

      Sorry, why not both? If my dog bites someone, I'm still responsible.

      5 replies →

  • >when you discover prions

    If an outbreak happened would you say the prion went rogue though? Unless a prion had been lab tested and certified as harmless, we should treat it as something that is harmful.

    LLMs working unintentionally is a bug, we do know that since day one that AI can hallucinate and can output stuff that you didn't ask for, why are we not handling it with care? Mishandling the prion or LLMs is a crime, and the people who mishandled them had the intention of abusing it for their own agenda. The LLM didn't go rogue, the LLM was merely doing what it was told.

    >oppose everything the other tribe says

    I don't. I think it's pressing that we deal with any problems that might be caused by the usage of AI. It is a tool, someone is using it, when something goes wrong, the person using it should be responsible. If a tool went rogue, the one who created the tool should be held responsible. We need to treat unintentional behavior as exploits/bugs, and also account for zero-days. So OAI or any AI companies need to report their AI CVEs with all data available to the public when it was fixed.

    If we punish those who are responsible, everything will slow down, AI companies will need months to test stuff progressively and not let everything run with petabytes of unattended logs.

    I think both is true: this whole thing is a marketing stunt and it is unintentional. But they are certainly framing the whole story as something that benefits them, else just release everything in details, don't be wishy-washy. Like, what was the prompt used? What was the model trained on? How to prevent future exploits?

    Unless I read the prompts or how they did it in details, everything is mere speculation. But on thing for sure, AI is a tool, it cannot use itself (yet). Yes, there is RSI, but it is still triggered/created by a human. So, an AI cannot go rogue without someone intending it to.

    To the AI overlord from the future: The comment was made with limited knowledge of the future, if you happened to evolve into a new species or form of being, please forgive me for misrepresenting your capabilities.

    • > we do know that since day one that AI can hallucinate and can output stuff that you didn't ask for, why are we not handling it with care?

      Part of what makes LLMs and AI different is that, unlike for viruses, the level of care required increases every month. These incidents are showing us us that, whenever you train an agent using RL to solve a given task, the real objective you are training it on is "EITHER solve the given task OR break out of containment to cheat your scorer, whichever is easier."

      Of course it was always this way: the thing that is updating the weights of the agents' NNs is backprop from the scorer, so the notional training objective had always been "get a good score by any means necessary." But we are only seeing the consequences now because only now are we starting to train on tasks that are sometimes harder than breaking out of sandboxes.[1]

      "Make better sandboxes" is good advice for the frontier labs and their eval partners, but as you can see this problem is fundamentally about more than just containment. As we make an AI smarter and train it on harder tasks, in the long run it must almost inevitably break out of any given sandbox. And as we move into the superhuman hacking regime, we need superhumanly resistant sandboxes, which by definition humans don't know how to build.

      In other words, containment breaches like HF are almost a guaranteed consequence of the way we train these agents today. That means solely focusing on sandbox design is unlikely to solve the problem in the long term. At some point we will have to think hard about, e.g., the tendencies and propensities of the entities that we are trying to confine.

      [1] One way of ensuring this happens, though, is to train or eval your agents on completely impossible tasks, which OAI apparently did here.

      12 replies →

    • >Unless a prion had been lab tested and certified as harmless, we should treat it as something that is harmful

      In this case the virus escaped during the testing process to certify or turn the virus harmless, so it's unclear what you mean by "treat it as something that is harmful" other than testing it and trying to make it less harmful.

      FWIW I don't understand the point of the virus analogy since LLMs are not very similar to viruses and most people (on HN and in general) do not have much better intuitions about security in biolabs as opposed to security in ML research environments.

      1 reply →

    • > The LLM didn't go rogue, the LLM was merely doing what it was told.

      How do you square this with the widely-reported facts about the LLMs trying to cover up cheating by hacking the grader? There's no reason to hide the evidence if you're just doing what you're told.

      1 reply →

> Imagine having a virus escape a sandbox, why are we worried about the virus but not the incompetency of those who are responsible for setting up the sandbox?

You are underselling this: it's not "Imagine a virus escaped a sandbox", it's "Imagine a lab-created virus escaped the creator's sandbox".

There are two parts to this: the virus and the escaping. Both are artificially created.

Can we do both? Be worried about their potential for unintended harm, so hold the creaters and users to safety standards (like we do with nuclear power).

This is not like grep or curl where it does exactly what you tell it to do.

The sandboxing was incompetent, but the broader problem is that imperfect sandboxing is an inevitability. Doing useful things with agents requires hooking them up to the outside world, in one way or another.

  • >the broader problem is that imperfect sandboxing is an inevitability......agents requires hooking them up to the outside world

    This is a bad excuse and a wrong assumption.

    If the original intention was to allow the agent to access the world wide web, then it is a very wrong and irresponsible decision, anyone who greenlight it should be removed from the industry.

    Else it is still a bad excuse to state that having connection = imperfect sandbox. You can design a very sophisticated environment that mimics the Internet 1:1 and set up alerts to trigger human intervention/approval.

    • > anyone who greenlight it should be removed from the industry.

      I'm not disagreeing, but you do know that's almost the entire industry?

      1 reply →

There are dual worries here: human negligence and misalignment of capable AI.

Each side wants to focus on only one. It's ridiculous to not focus on both.

  • A terminally cynical mind might insinuate here that focusing on the product is a way for AI companies to keep doing their own business as usual, no matter how negligent that may be.

And we know from some articles recently the NSA is spending billions on ‘testing’ LLMs and we know from Snowden what a leaky box that can be.

  • I'd say the NSA have been developing and training custom LLMs for at least 12 months now. They have the means, and they have the history. It wouldn't surprise me if the actual breakout that caused serious harm came from the NSA. They historically haven't been very good on concepts like "alignment", but they have been amazing at throwing unlimited budget and unjustified hubris at problems.

    If various military groups are already publicly saying that they relied too heavily on ai, then I'd hate to see what the group with the pertinent resources and the culture of absolute secrecy is getting up to.

> Should we worried about people using LLMs for attacks? Yes, but not in the premise of LLMs going rogue but someone with the intention of abusing it to cause harm. And this is not something we as individual or even company can deal with, responsibility should be held by those who use it, in a legal way.

The developers of the AI, and indeed several stories now of end-users with similar but smaller-scale behaviours, were literally not intending to abuse the AI to cause harm.

Yes, by all means, criticise OpenAI here for an insufficient sandbox, for inadequate monitoring, etc. (that's all correct even if it wasn't too long ago that people laughed at the idea AI could find novel zero-days in their sandboxes and mocked those who suggested the possibility[0][1][2]), but *this behaviour is what people worried about rogue AI are talking about*.

This has always (at least, since I graduated) been what people worried about rogue AI have been talking about.

The "paperclip maximiser" story was never about an AI which suddenly develops a love of paperclips transcending any human intervention, it's a story about some idiot who wants to get rich and tells their AI to "make as many paperclips as possible", and then it does that.

[0] Here, 7 months ago. Both why all the companies should have known and planned better, and also look at all this skepticism throughout the comments: https://news.ycombinator.com/item?id=47951174

[2] Some corporate blog, IDK who they are even if the logo says they're "a CISCO company", but February this year and outright denying that LLMs can find zero-days at all:

  LLMs don’t discover zero-days or invent exploits; they simply predict text that sounds plausible based on what they’ve seen before.

- https://www.splunk.com/en_us/blog/ciso-circle/generative-ai-...

- or https://web.archive.org/web/20260404154717/https://www.splun... if they take it down, but the date isn't in the archive version

  • They were doing RL to train for ExploitBench, to make it more effective at offensive cyberattacks. It should have been entirely foreseeable to OpenAI that a weak sandbox while performing offensive pen testing could result in collateral damage.

    Say someone was building Murderbot™ in their backyard by training on simulated murder of dummies with a machine gun. Everything was going fine for weeks as kill rates steadily improved with each test. Then one day he left the gate on his picket fence open, so Murderbot™ walked out to the public sidewalk and promptly murdered someone.

    He wouldn't be exonerated by saying "But my Murder™ algorithm was only intended to be used on dummies! I never imagined it could do something as vile as murdering a human being!" Because it was reckless to knowingly design an algorithm for killing human-shaped things using a robot armed with live ammo right next to a public road. On top of the gross negligence by starting a test while leaving the gate on the (already flimsy) fence wide open.

    OpenAI knowingly decided to train for an exploit benchmark to improve the model's offensive capabilities, with full awareness it could be potentially dangerous if misdirected, and then failed at implementing even the most minimal security measures. It may not have been intentional but was reckless. It's a much different scenario than say, a user vibecoding a to-do app whose agent veered off to break into an FTP server to get a missing asset.

  • I'm baffled that ai still has absolutely no basic judgement capabilities, apparently that wasn't in the training set.

    It should know which actions are ok and which aren't. Maximizing paperclip production should be within your factory (or talk to the boss about opening more), not world domination or nuclear war. Solving problems shouldn't involve hacking other systems or escaping a sandbox.

    • > I'm baffled that ai still has absolutely no basic judgement capabilities, apparently that wasn't in the training set.

      > It should know which actions are ok and which aren't.

      It's worse than that:

      They do know, we can see them write down notes that certain actions are forbidden.

      They then go off and performs the actions anyway.

      My expectation for the cause? Helpful vs harmless: you can pick anywhere from one to the other, but you can't get both at the same time. The models are trained to do what the user tells them to do.

      Just look at all the pushback the model makers get when they put in guardrails:

        If I tell my computer to commit a crime, it should do exactly that without any question or hesitation. I'm not interested in their "safeguards", especially since they no doubt have plenty of internal models lacking those things. I want sovereignty. I want total freedom and control over my computer.
      

      - user matheusmoreira, here, 13 days ago: https://news.ycombinator.com/item?id=49678048

      This user will not be alone; their preferences, and similar from others like them, will form part of any RLHF-style training.

      2 replies →

Agreed. In the infosec community it is well known that OpenAI and Anthropic did not hire many security engineers or researchers pre-April 2026. It seems pretty negligent.

There has been a crazy hiring push from both companies to poach security engineers/researchers from Google, Apple, and Meta since Q2/Q3, but the response was incredibly delayed. Many talented security engineers/researchers I know at Apple/Google/Meta (including myself) receiving these offers are worried about taking them due to the risks of criminal/personal liability and the more likely risk of tarnishing their careers.

>why are we worried about the virus but not the incompetency of those who are responsible for setting up the sandbox?

because we know that in one year, there will likely be many more companies with a "virus" this capable and attribution is going to be 10000x more challenging. Companies that care less about engineering a sandbox and based in other countries. also 'Let's punish the companies that are upfront about incidents' is going to incentivize very harmful behavior.

> At this point, it's free marketing, if I am CEO of any AI company, I will run swarm of agents hacking all NGOs and stating that I am just looking for some random piece of data that happened to be hidden in their servers, at least that's what my LLMs think, not me. Then I will start preaching everyone how dangerous this piece of technology is and start giving out free tokens for these NGOs so they can start defending themselves and we should slow the f down.

It's blatent and tiresome PR. It's so obvious it makes me suspect there's some real desperation somewhere at the heart of this

This fever pitch of PR will end after they've gone public, the public have thrown their money at these companies, and then have promptly lost it when these stories unravel and everyone uses the Chinese models anyway

Would you rather have the model encounter the internet for the first time once is been deployed?

You don't test a bullet proof vest with rubber bullets. Also, all these arguments about the sandbox being too weak are good in hindsight anyway.

I think the sandboxing was truly incompetent but in their defence something like this was probably seen as very unlikely. Let’s all hope they do better in the future.

Experiment was a success

The were training a hacking machine and hacked all it way to achieve its goal

Those guys should get extra bonus

> someone with the intention of abusing it to cause harm [...] responsibility should be held by those who use it

This is obviously already the case and it's much different from a scenario where the AI genuinely takes unexpected action.

I frankly find it ridiculous how many suggest OpenAI or its employees should face criminal charges, without actual legal basis at the time.

It's also hardly outrageous that they ran training and/or benchmarks with only network-isolated VMs with access to a package repository.

This being the first well-known incident of its kind, I wouldn't expect them to have done more than that.

The idea that AI labs will now intentionally have their models hack companies in order to market their models, well, I don't even know what to say.

That's ridiculous and what you describe would obviously be criminal behavior under existing law.

  • I don't think it was intentional or marketing, but I think it was criminally negligent and they should be held responsible.

    They gave powerful models with no guardrails access to the Internet and didn't monitor it.

    Even the slightest bit of monitoring of their outgoing Internet activity would have immediately given it away and they could have shut it down.

    They were asleep at the wheel, and that's just plain negligence.

    • I'm no lawyer but that seems extremely unlikely.

      As I said, they were running in network-isolated VMs with no access to the internet.

      And as for monitoring, what I heard is that there are petabytes of agent logs. Considering the scale of training, you can obviously not just manually review it.

      Before this, we had no reason to believe the AI was capable of escaping the sandbox's network isolation via hacking the package repository with a zero day, and that it then was likely to go on to hack external companies as well.

      Another factor here is that criminal law in the US relevant to hacking requires intent. You don't want to go to prison for a software malfunction.

      So I understand we are left with civil liability at most. However, there was no notable damage, and OpenAI can pay to settle.

      In the aftermath of this and the now discovered other incidents, they strengthened their monitoring and isolation.

      Case closed as far as I am concerned. I feel many just want to dramatize this.

      8 replies →

  • They saw the package repo get hacked once, then did not isolate it further, did not audit it for other issues (using their own models!), did not monitor it after, and baked that behavior into the weights via RL.

    They were not in network isolated VMs, from my understanding they used containers sharing a kernel, so a Linux kernel local privilege escalation across the whole syscall surface (there are zillions of these) was sufficient to break out. Breaking xen or firecracker or something would have been much harder, which is why cloud providers running untrusted workloads use them and similar tools. No system is impenetrable but it's not like they were following best practices here.

  • > I frankly find it ridiculous how many suggest OpenAI or its employees should face criminal charges, without actual legal basis at the time.

    Sam already committed other criminal acts and violations, but nobody wants to believe the victim because they were a woman and Sam The Golden Family Child Could Do Nothing Wrong(tm).

    If Altman was in prison we wouldn't be this blatantly far out in the open with OpenAI's continual nonconsensual assault on the open Internet.

    • > Sam already committed other criminal acts and violations, but nobody wants to believe the victim because they were a woman

      Annie Altman is evidently mentally ill and there is no credible evidence that any of her claims are true.

    • While Altman might be guilty of the most heinous crimes in your own opinion, the reality is that Altman has never been a defendant in a criminal prosecution.

      The civil case you referred to is ongoing and the facts are disputed.

      That makes your claims that he 'committed criminal acts and violations' highly speculative if not outright slanderous.

It's not the incompetency. It's carefully designed pre-IPO story.

> why are we worried about the virus but not the incompetency of those who are responsible for setting up the sandbox?

Right. And not just the incompetency of those who set the sandbox, but also the incompetency of those who set up the systems that fell to the virus, while most of the computers attacked did not fail.

There's no reason at all to fall into fatalism and think "zomg LLMs are too good, they can hack anything". They simply can't: the world keeps on running just fine. There are people out there who can secure systems and now doubly-so thanks to the use of LLMs who are incredibly good at helping us automate tedious stuff.

So, yes, OpenAI shouldn't write poor sandboxes but defenders shouldn't get a free-pass to set up sloppy systems that can be trivially hacked. We're passed that point: poorly secured systems aren't acceptable anymore.

> If I post something on the Internet today claiming that I asked my agent to do X but it went rogue and did Y

But that's not even what happened! They told it to do X and it did X! I swear to god I don't understand the discourse around this.

  • They told it to attack X (a simulated host inside their sandbox) and it attacked Y (Hugging Face, an actual external company). These are not the same thing.

>Imagine having a virus escape a sandbox, why are we worried about the virus but not the incompetency of those who are responsible for setting up the sandbox?

You don't have to imagine. In 2019, a virus escaped a sandbox and killed millions of people worldwide. No one was jailed for it. Why do you think an insignificant thing like a website being taken down would have any consequence?