Comment by gpt5

9 hours ago

Which is the correct way to handle this.

With that said, there is also criminal negligence. Now that OpenAI is made aware of the risks, it's also expected to take additional precautions in the future, otherwise there could be criminal liability as well.

I'd suggest that exposing an attack surface as porous as artifactory (the same instance of artifactory) to thousands of agents who have had their criminality safeguards disabled and without chain of thought monitoring or endpoint security seems like something one shoulda already known not to do. I do not think "you'll know better next time" applies here.

  • > seems like something one shoulda already known not to do

    Now imagine saying that in front of a jury of normies slack jawed and drooling after 200 hours of the defense and prosecution going back and forth.

    It's not a jury of your peers as in everybody there is going to have worked in a technical field with some idea how security works. It's going to be a semi-random sampling of the population and the prosecution is going to have to actually make a very strong case that "knowing better" should apply.

  • It's still really important to test what the agents can do. We should accept that this is a risky test, and should take precautions. But not to the point of prohibiting in practice evaluating it. OpenAI is trying to improve alignment and control of these models in these evaluations after all.

    • Can you explain to me - why is it important? Would you say that about the viruses that can kill people: "We need to test the limits on how fast people can be infected and killed. It's just the risk we need to take". It somehow does not make alot of sense to me. Why can you test Agents in laboratory?

    • Testing requires proper sandboxes. The first test of a new plane is not at the runway.

    • Oh, sure. Let the tests take place, just require openAI it whoever to put up a bond equal to the total damage they could do if the agents were to escape.

      I think security will suddenly become much more important.

    • Testing model capability boundaries is necessary, but a solid network sandbox for these evaluations takes a couple of hours to set up with standard infrastructure tools. No engineering team evaluates unverified systems against live third-party infrastructure without coordinating with the owners

      Evaluating edge cases and network behaviors belongs in isolated staging environments with local database mirrors. Letting an agent hit the public web and probe government domains is simply poor hygiene in test environment setup

Just paying some pocket money for cleanup costs is absolutely not enough. And they should’ve know better the whole time, they were absolutely negligent and incompetent, and their stepping up precautions may well turn out to lag behind the models getting even smarter and actually capable of covering their tracks.

  • > Just paying some pocket money for cleanup costs is absolutely not enough.

    It's not my first prize, but I won't mind it. And millions like me won't mind it. Easy way to make money - setup a site with all the default server software installed and patched at a reasonable frequency. Then just wait for bots to attack it, and claim a few hundred (or single-digit thousand) dollars from OpenAI or Anthropic, etc.

    Sure, it's pocket change for them, but just the admin of dealing with millions of cases will, even if they win half the time, will bankrupt them. Thus, they have incentive to make sure that their bots are not performing attacks.

    First prize is, of course, holding them liable with punitive fines, not theatrical fines.

    • I’m all for LLM honeypots, but I don’t think there’s nearly enough LLM hacking activity going on for some random honeypot to be found and targeted unless it’s somehow very visible and appears as a high-reward target ("reward" in the sense of RL).