Comment by hgoel
8 hours ago
The HuggingFace incident was the culmination of OAI allowing thousands of agents of various different models - with no clarity on which stages of development they were at (for all we know, some of those models did not have safeguards trained in yet) - to run for at least many weeks without any monitoring in place and with very little thought given to the warning signs (all of the various messageboards) before the incident happened.
Theirs was an example of the "reckless waste of resources" I mentioned.
We are apparently supposed to believe that OAI takes this incident so seriously as to seek regulation after they have been found to be hiding most of the details of the HuggingFace hack, limiting what their so-called third party investigators can see, and on top of that, had no concerns when they rushed to spin up a 10,000 agent swarm of an internal model, running for several days, to try to get ahead of researchers rumored to have made meaningful progress on a well known mathematics problem.
Edit: Actually, we were explicitly told that some of the models used had safeguards relaxed!
'Model-level safeguards were reduced by design. OpenAI said that "deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities"'
https://en.wikipedia.org/wiki/2026_OpenAI_agent_cyberattacks...
There is nothing that could prevent a bad actor from replicating exactly the same thing with the given goal of e.g. gaining control of critical infrastructure or extorting money. Except for maybe economics.
Same can be said about a hundred other things in the world. All the way from knives to nuclear.
Bad actors could and will train their own models eventually. So what's the point of crippling frontier? It will only delay preparations for dynamic of new world prolonging the fake sense of relative safety and temporarily lowering motivation to find actual robust mitigations.
Letting bad actors dictate the pace of technological development is certainly one option, but not a good one.
There's nothing stopping anyone from doing it, even without AI. People have proved entirely capable of doing a lot more hacking than happened here.
What company, product, or period of industrial history do you think met your standard of prudence?
What are you trying to say?
I'm asking you a question. What is an example company or industry that meets your standards of prudence? For me it would be, say, Swagelok. What is yours?
1 reply →