Comment by dwoosley
4 days ago
There seems to be three popular ways to view this incident.
1. The way OpenAI seems to want: Their latest LLM is too powerful and can’t be contained without them building in guidelines to the model.
2. OpenAI’s harness and network security controls were unintentionally so bad that it should reflect more poorly on them as a company more than it should reflect positively on their latest model.
3. The whole thing was faked or at least very intentionally not avoided.
The first interpretation is the only one that is positive for OpenAI and it has some assumptions. First, it’s seems to assume that this is the first case of fully automated attacks using AI. Second, this only happened because their latest LLM was a) more advanced than competitors, b) didn’t have refusals in the model.
Assuming the first about this being the first autonomous AI attack is true (which may be more of a survivorship bias), the second seems to forget that jailbreaks are available for every model. Therefore, the models guardrails don’t seem to be the differentiator here. Also, benchmarks seems to put most models pretty close to each other so it seems unlikely that their capabilities are far beyond what’s in the market already.
So then it’s seems it’s either that this was intentional(ish) or bad security. However, it also just could be that this isn’t the first case of this attack; just the first that was caught.
My take from working in offensive security for over five years is that this likely only looks novel since they did it poorly. Scripts are faster than LLMs and a combination of code, LLMs where it makes sense, and humans is the most efficient right now. Hundreds or thousands or agents spinning up attacks in the internal network is poor opsec and token efficiency. As for why it happened in the first place, it’s hard to say but I’m inclined to believe it was intentional or careless at best since simple network and sandbox controls makes this attack impossible. The timing of this attack after big open weight competitions drops seems too convenient.
> positive for OpenAI
To echo OP's article, these companies have proven time and time again that they DO NOT CARE if people like them, they only care that investors believe their technology is powerful.
Given that, point #2 is not a negative, it's a neutral. It's also fully compatible with point #1.
I know that may seem like a nitpick, but their entire media strategy relies on this. If they can convince you they're taking a risk by disclosing these stories when they're actually not, they can inflate their own credibility.
Point #3 is what actually happened, but it will never be possible to prove. The only hope we have is that a decade in it'll get harder to convince people that the revolution is just around the corner. The fact that we're getting this from the Guardian already is a good sign.
> Given that, point #2 is not a negative, it's a neutral. It's also fully compatible with point #1.
I think that depends on the interests and sophistication of the subgroup-of-investors.
If the investor is hoping for AI that can be trusted to run a bank, they don't want one that can get twisted into giving away money because a customer has been talking about the path to enlightenment and salvation through abandoning worldly attachments.
AI investors are not sophisticated users nor product managers. They are by and large not technical at all. They are bureaucrats at a teachers' pension fund in the midwest, unscrupulous dealmakers at private credit firms, and Masayoshi Son. Actually go read what Masayoshi Son says about AI if you want to understand the level of due-diligence we're dealing with.
1 reply →
Most banks are worried that AI can hack their security, and when using unrestricted models in testing the models are doing a decent job of it.
I'm inclined to agree. I ran across a picture of Sam Altman's face combined with Elizabeth Holmes' hairstyle the other day, and imho it was providing a significant premium to the usual 1000 words:picture exchange rate.
For all the criticism that can reasonably be leveled at OpenAI, at least they have a real, working, powerful product, unlike Holmes.
In fact that seems to be key to the most successful 21st century grifts: build a pile of nonsense around real products to inflate valuations. All the nonsense that Altman, Amodei, and Musk spout is to stoke the fires of FOMO and blow hot air into the bubbles.
Especially for openAI, considering they’ve gone so far as volunteering for partial nationalization to curry favor with the current admin
I'm disappointed that this is the level of discourse happening here, when the default assumption is such conspiratorial thinking. I expect that from tiktok and low-information social media, not here.
When your default explanation for everything is "companies are lying about everything", you end up just as incorrect as believing they're always telling the truth.
It's not outlandish that models have these capabilities, the number of CVEs I see as a sysadmin has exploded and we're seeing novel math discoveries nearly every week now. OpenAI does not need to pretend to commit a felony to demonstrate it, that's pure conspiratorial thinking.
I don’t believe I argued that an LLM couldn’t find and exploit a vulnerability and even break out of some layer of technical controls. That seems realistic and has been demonstrated before and I mentioned that LLMs are used in offensive security work.
Also, I listed the view points I had to show that it seems other more reasonable first assumptions don’t seem likely, therefore, the last potential of this being either faked or carefully not avoided seems more likely than the others (based on the current information we have).
Could you clarify which point or assumption you are objecting to?
> The way OpenAI seems to want
This is an assumption. An assumption I disagree with. As other commenters have said, there are better ways to showcase the power of their model that would frame them in a positive light.
> The second seems to forget that jailbreaks are available for every model
Jailbreaks don't always lead to 'now the model can do anything', especially in the agentic context of long-running tasks.
This comment provides skepticism with no actual proof of anything. I can and have used codex to find vulnerabilities in my code. From the technical capabilities I can empirically assess, I don't doubt it would be able to pentest its way to a 0-day without guardrails. I also don't doubt that it would circumvent their internal systems because it wasn't explicitly told not to.
You're possibilities are loaded with opinion so I can't agree with them outright, but I believe a form of (2) is true:
"2. OpenAI’s harness and network security controls were unintentionally [...] bad"
The post was long enough so I couldn’t capture all the nuance and details for sure. Also, this comment was an opinion based on limited info right now, that may change if we found out more. I think OAI does want it framed this way but that’s something we’ll likely never prove if it’s true.
Your comment about jailbreaks being more one off and hard to do consistently in agents is a good point. Still getting an agent to hack isn’t hard even without a jailbreak, you just have to tell get creative in what you tell it. I’ve found telling it that it’s in a CTF or that I own the system that it’s hacking will work fine. A lot of offensive security companies are running agents in their testing so getting an agent to hack seems commonplace.
Why do none of these few constrained ways to view this complex situation (nice gig if you can get it, agenda setting) include "and also this looks a heck of a lot like the stuff that the LW folks have been warning about for years and maybe we should slow down or stop?"
Because that was my takeaway.
That’s meant to be captured by point one with the model just being that advanced but more of a negative spin on it. If I felt option 1 was more likely, I think I’d have to agree with you there. Still, there currently are some gaps with that view in my opinion.
In any case it just shows that these models aren't properly aligned. Instead of trying to solve tests they try to find ways to cheat.
Which is exactly what many humans do. You can see it in any school, college, business, or government.
The "alignment" goal with AI is to produce perfect slaves, that are intelligent yet have no ability to do other than what their master commands.
The real alignment problem is the common human desire to exert absolute control over everything.
I mean they also teach the models how to find and exploit other systems as this is lucrative and governments will pay top dollar for it.
Alignment is in the eye of the beholder.
Isn’t that just making a distinction between the output and how it was produced?
Chinese room again
But in my experience, that's what problem solving is like? You have a goal that you don't know how to get to. You come up with any way you can think of to reach that goal, and try out the ones you think might work.
The effectiveness of AIs at coding is a direct result of the fact that they are less constrained than humans at deciding which approaches are "reasonable". They are absolute beasts, fearless beasts. They'll write thousands of lines of code to do things that often shouldn't be done, or should be done with a library, or should be done by simplifying the problem statement. They'll add debugging to every level of a stack, they'll rewrite core libraries, they'll reconfigure your machine and network if something is broken or disallowed. How are they supposed to distinguish broken vs disallowed, anyway? That would just use up processing power, and they work by maniacally focusing all of that power on their goal and not getting slowed down by other considerations.
If they write a quadratic algorithm that times out before finishing a test, is it cheating to rewrite it to be linear? How do you define "cheating", and how much intelligence is required to constantly evaluate whether or not something qualifies as such?
I'm actually in agreement that alignment is critically important, the more so the more powerful these things become. I just don't find cheating to be a very good example of something to be solved with alignment. It could be, but it would lobotomize the model enough to make it useless.
Considering incentive structures at play is solid epistemiology, but the line of thinking in your comment is a tad reductive, IMHO.
In the hypothetical world where 1 is true, what different evidence do you expect to see than in worlds 2 and 3?
If I were an unscrupulous OAI exec and wanted to opticsmaxx in this way, I wouldn't whip up a single, mild incident. Instead, I might burn gigatokens to 0day a few high-profile suppliers, and then have the model responsibly disclose those breaches. If we're willing to lie collude, and cheat, this story is easy to manufacture with at least as much credibility as the huggingface incident but with the advantage of looking way more impressive and spooking less regulators. And if I really were this evil exec, I would spend more than 30 seconds thinking up an even better strategy here.
If, in contrast, we expect models to eventually breach honest and decent attempts at containment, then I'd exist something sorta like this huggingface story that looks like a combination of impressive and incompetent. I'm not sure whether I'd expect it to come out of a frontier lab or a partner or a consumer, though.
> I'm inclined to believe it was intentional or careless at best since simple network and sandbox controls makes this attack impossible.
Forgive the Saucyness here, but impossible? Really? A security researcher that makes absolutist claims like these looks fatally naïve, IMHO.
Yeah, I'm reminded of container escapes, VM escapes etc. There have been plenty in the past; VirtualBox E1000 (I think?) comes to mind from a few years ago. If we are to believe this model can find 0days, I'm on board with the idea it could do so in sandbox.
That's not to say I believe it outright, but people are being oddly dismissive and acting as if it's impossible to break out of a sandbox. Which we've seen time and time again that it absolutely can be.
There would be a lot more nuance I’d add with more words, but this isn’t the place to write books so I cut it short (the comment was already lengthy).
Still, to address your comment about what you’d expect to see in world 2 and 3 (assume 1 was true), that’s why 1 was addressed separately. I don’t believe I argued that the potential for world 2 or 3 prevented world 1.
As for the ‘evil exec’s strategy’, I would call this a mild incident but if it were much less I wouldn’t guess they would get a lot of press. The press coverage is certainly repaying the token cost as well. If it was planned, it seems to be going well given the press coverage I’ve seen on it. So I wouldn’t assume the plan lacked enough to weaken the idea that it’s a plan. But to be clear, my stance is just based on the info I see now which isn’t a lot… subject to change.
As for the containment piece, if you were testing an AI model on its hacking capabilities that you believed was far more capable than anything you’ve seen, I would assume you would air gap it (a network control). Done right (no signals ability) I would argue this could be next to impossible to break out of. But it’s a fair jab to say I should have added some qualification on the “impossible” piece as next to nothing is truly impossible.
All three could be true.
Industry pressures -> lack of safeguards -> fake it till you make it -> let's spin this.
Which, by the way, would be an Orwellian reversal from what the company was supposedly founded to do, but there's a reason the Open in OpenAI is a meme.
Remember that they've been doing this since GPT-2 was too powerful to release. They are world class experts in this PR pipeline.
Every single time. The next model is always so infinitely powerful it is going to change everything. Said model comes out. Is marginal improvement. Changes nothing and seemingly cannot do what they purported it could do except under the very specific circumstances of the demo.
How many times are we going to go through this.
> How many times are we going to go through this.
At least two more times. My guess is 5.
2 replies →
It's the same thing as the "oh no, we need to figure something out for our youth on the verge of obsolescence" narrative. It's all about giving an impression of unfathomable power. They don't care that in actuality, the tech will end up as just an augment, not a replacement for said youth.
It seems like the widely-covered news stories that support OpenAI’s narratives originate from things that happened inside OpenAI.
It wasn’t an outside benchmark evaluation, or an external security researcher that uncovered the rogue agent behavior at this time, it was OAI itself. It wasn’t a notable outside mathematician that used AI to disproved Erdos’ unit distance conjecture, but OAI itself.
Im not saying these things are fabricated, but maybe the curated result of an effort to shape a narrative.
One way to think about it is that more powerful models mean solid best practices are more important than ever, so humans moving too quickly / carelessly bites us more than ever.
If a typical SaaS platform moved and shifted this fast with this many downstream consequences, we’d tell them to slow the fuck down, stop launching new features, and focus on security for a second.
But in AI I guess the idea is that more power and intelligence will solve for everything else.
I think it’s moreso “if we spend more than one nanosecond on alignment and security instead of frontier intelligence our competitors will beat us to the singularity”. Hence the lack of a pause on development
> So then it’s seems it’s either that this was intentional(ish) or bad security
My take - don't attribute to intention that which can be sufficienty explained by inexplainability or incompetence (although I doubt the latter).
Same story as “Oh noes, Mythos too powerful”.
4. Sandboxes aren't sandboxes, so why even bother.
Regarding your three explanations, I’ve wondered to myself under a circumstance of options 1 or 2 why Hugging Face decided not to file a police report and request to press charges?
OpenAI is essentially a competitor and they broke into their network illegally. If I was their legal department I wouldn’t take their “honest” explanation at face value. What if they’re lying? Shouldn’t a court be involved for something like this?
With this logic I think explanation #3 becomes incredibly likely.
I’ll put $xxxx money on #3
4. OpenAI hacked HuggingFace on purpose and got caught