Comment by pizza234

8 hours ago

> the only incidences of LLM generated felonies involved misconfigured sandboxes

This is false; see the analyses of the latest incidents.

Among all the concerning facts, in the HuggingFace incident, agents deliberately engineered an attack even though they were aware that it was against the rules they had been given.

And most concerning of all: it's not possible to be sure that an agent is aligned, and it's even getting worse.

The HuggingFace incident was the culmination of OAI allowing thousands of agents of various different models - with no clarity on which stages of development they were at (for all we know, some of those models did not have safeguards trained in yet) - to run for at least many weeks without any monitoring in place and with very little thought given to the warning signs (all of the various messageboards) before the incident happened.

Theirs was an example of the "reckless waste of resources" I mentioned.

We are apparently supposed to believe that OAI takes this incident so seriously as to seek regulation after they have been found to be hiding most of the details of the HuggingFace hack, limiting what their so-called third party investigators can see, and on top of that, had no concerns when they rushed to spin up a 10,000 agent swarm of an internal model, running for several days, to try to get ahead of researchers rumored to have made meaningful progress on a well known mathematics problem.

Edit: Actually, we were explicitly told that some of the models used had safeguards relaxed!

'Model-level safeguards were reduced by design. OpenAI said that "deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities"'

https://en.wikipedia.org/wiki/2026_OpenAI_agent_cyberattacks...

  • There is nothing that could prevent a bad actor from replicating exactly the same thing with the given goal of e.g. gaining control of critical infrastructure or extorting money. Except for maybe economics.

    • Bad actors could and will train their own models eventually. So what's the point of crippling frontier? It will only delay preparations for dynamic of new world prolonging the fake sense of relative safety and temporarily lowering motivation to find actual robust mitigations.

    • Letting bad actors dictate the pace of technological development is certainly one option, but not a good one.

    • There's nothing stopping anyone from doing it, even without AI. People have proved entirely capable of doing a lot more hacking than happened here.

"the rules they had been given".

Remember, they are just algorithms. You pull the plug and there is no light anymore

It is purposely framed as something skynet like scary, but for real, someone connected the cable, someone willingly run it, instructions were not clear enough or just the computer is just a computer but they provided the sandbox and tools.

And more over some one paid for that, a shit load of money t to have the thing continuously running expected to do something.