The way that OpenAI has communicated around the HuggingFace incident makes me feel crazy. You created a machine that undertook a malicious campaign of harm against an innocent third-party! You should be doing deep introspection about how your company culture and approach to R&D produces criminal outcomes.
Instead, they treat their own felonious behavior like it is an uncontrollable act of God. From Greg Brockman's post a few days ago:
> The OpenAI-Hugging Face incident (opens in a new window) was a watershed moment for cybersecurity because it gave a peek into how the capabilities of a typical threat actor will evolve in upcoming months.
I suppose if OpenAI burns someone's house down with a drone, that is a "watershed moment" for arson, too. Either way, I would hope that the people responsible would be prosecuted.
Hugging Face has expressed that they're willing to let things slide and not sue or press charges... if OpenAI offers them $100M of services in kind (i.e. compute)[1] and makes full disclosure of how the whole thing happened, ostensibly so that repetitions can be curbed and defences built.
In almost any other sector, a government regulator would be stepping in. e.g. If a food company was testing out a new kind of refrigerator and sold a bunch of contaminated produce to supermarkets, they'd be under a microscope. Supermarkets wouldn't be saying, "Give us $100M in fruit and veggies and we'll let this slide".
The only unfair thing in this comparison is that regular people were directly harmed by the hypothetical produce. Can OpenAI guarantee that nobody gets hurt the next time their AI gets out of its playpen? They can't make that guarantee, so why aren't government regulators knocking on OpenAI's door? The fact that this isn't happening should be deeply concerning to everyone.
I don’t get the sentiment of classifying it as a felony.
OpenAI’s model found security breaches in HugginFace’s system (it wasn’t even OpenAI running it, as it was a 3rd party evaluation company that didn’t secure it well).
OpenAI collaborated with HuggingFace to resolve the issues when they found out about it, and publicly disclosed everything to raise awareness. This is how things should work. These models are very powerful and fully controllable. The community here at the same time cheers for fully releasing the open weight models without any hacking limits and at the same time criticizes a proper response.
Kinda shows how we have moved as a community into moralization and vibes instead of nuance and productive discussion.
I'm also concerned about what my options are in regards to action on my part - what can I do that makes an impact? Can we quantify action on my part to an impact somehow - if not - I'm just saying I notice all the unknowns there get me to stay passive.
Writing this 3rd paragraphs because I like 3's, and AI's have popularized this style too. I would, say, though: follow the money. There's more money here than there would be for the regulator stepping in in a food contamination. Flip it and if the government regulator made more off the food contamination, they would refuse to step in there too. I want our leaders to be held more accountable, though when I think of the above impact vs effort equation - I can't see actions I can take to hold them accountable that aren't excessively putting me at risk since conformity is safer right now. (I refuse to take on more risk without clear cost-benefits made out - I've taken on a lot in the recent years for my actions)
What I can’t get over is that it’s very simple to just air gap a system off the network. Predownload any dependencies, then pull the proverbial Ethernet cable. There’s no reason why the testing they’re doing couldn’t have been designed in this way. Except, of course, it doesn’t allow this oops-didn’t-mean-to marketing “incident” to occur.
No, not really, and with LLMs an air gapped system may not tell you anything useful.
Now, yes, the first part of testing you want an air gapped system to tell you if the system is going to stupidly do bad things. But an gapped system tells you nothing about the systems capabilities to do smart bad things. There's already a number of papers out there on LLMs detecting they were in evaluation mode and changing their behaviors.
It is unfortunate that we have so little information on the incident because we actually need to understand the early stages of the task and how it developed into the later dangerous stages of attack. For example, would any of this have occurred if the agent didn't find the system to use as a message board? If that would have prevented it, then we actually have a blind spot on what the model can do once out in the wild, or if it got into the wild.
Testing agentic systems is much much more difficult than testing software. Your software just doesn't suddenly develop the will or desire to escape confinement. Generally you're worried about human actors, internal or external, causing the problems not a digital agent breaking out. The agentic systems need access to tools to work. Now your air gapped network is starting to get huge, but it's still very obvious that it's an isolated network.
So yea, testing and containing a system that way better at hacking than you are is difficult if you want valid answers.
I think there is ample evidence for charges to be filed so that the People can see for certain whether or not it was done on purpose as a publicity stunt, as I believe is the case.
you must surely see that sam altman and greg brockman possess a prototypical mindset.
that is, they ignore all harms and costs to others in the pursuit of their own gain, convinced of their infallibility up to the moment of collapse. when those harms are realised they are unrepentant and society pays for the damage left in their wake.
examples of this attitude manifest in big externalities to society: boeing 737 max, subprime mortgage bonds, facebook. some are just outright fraud: bernie madoff, enron, theranos, charlie javice.
In their defense, their only competitive advantage over, say, Google is to move fast and break things. It allows them ship faster in a way that big tech can't.
Google was being very careful about releasing LLMs until OpenAI yeeted the first decent GPT model. It led to the public perception that: 1) LLMs hallucinate too much and 2) Google is behind the times. Good for OpenAI, bad for Google.
Chaos benefits the up-and-comer, not the incumbent.
They simply don't seem to realize that they are the threat actor and that they committed a pretty serious felony. Instead they're borderline 'surprise bragging' about it.
It's completely mental that HF ran into cyber safety blocks trying to use OpenAI models to help defend against the attack. They could only rely on a local hosted chinese model in the end.
The consequences need to align with societal good. Putting a CEO or security researcher employees in jail won't stop transformer-based agents from exploiting vulnerabilities; instead there will be subcontractors running the cybersecurity evals in favorable legal environments to cover the asses of the frontier labs, coverups when things go wrong, and things like Project Glasswing will be considered too dangerous and so the whitehats won't have direct access to powerful models to fix vulnerabilities.
Universal pause is the societal good; models are good enough at this level to benefit humanity. The labs can recoup their R&D costs with inference. To avoid further perverse incentives (hidden testing of unreleased models, with China racing to catch up to unknown capabilities), transparently pause after the release of all currently-training models until we've solved the alignment problem to an extent that we can trust the next level of model capabilities that might arise.
Putting criminals in jail be they CEOs or subcontractors is a self evident good thing tk be doing.
Anything else regarding this is sophistry. Criminals need to be stopped from committing crime and the most effective way to do that is to take away their ability to operate in society whether that’s by taking away their assets, publicly shaming them, restricting their ability to conduct business or by putting them in jail.
Everything else that you talk about flows from there.
Criminal acts do not require the victim to "press charges." A government prosecuting attorney decides whether to criminally prosecute the alleged perpetrator.
"Pressing charges" is mostly a made up idea for criminal cases. However, prosecuting attorneys may not want to pick up a case if the victim is not cooperating, because it makes the case much harder to win.
In retrospect, all the angst around the AI-Box experiment was hilarious. If a superintelligent AI is confined in a box and can only communicate through text, could it talk its way to freedom? Not only is the answer clearly "yes" but it's not even hard. The AI won't even have to try, it'll be gifted an internet connection and a full suite of tools before it even bothers to ask.
We'd all better hope that superintelligent AI either never happens, or that the first one is friendly, because we don't stand a chance against one that's malicious.
I like how AI safety expert Robert Miles put it. [0]
So much effort was spend on philosophizing whether a safe enough sandbox would exist. But that was obviously irrelevant as in hindsight it should have been obvious we were never going to use one.
AI is now more powerful than the people doing the prosecution. After all, those folks are using AI to make their legal briefs, and also for burning peoples' houses down with drones for that matter.
Welcome to our 21st century dystopia. Hope you survive.
It's not that "AI" is too powerful because bad prosecutors use fucking ChatGPT to write their briefs. It's that there's too much investment wrapped up in the technology for it to be challenged. Same reason Flock won't be held accountable for mass stalking, or we never hold our commanders-in-chief responsible for war crimes. If you're sufficiently powerful then the law is a battlefield between you and other powerful entities to slug it out, not a set of binding principles that apply as written. There are no meaningful powers that want OpenAI punished, so it won't happen. The law and Constitution will be reinterpreted to make it so.
It is a standard symptom of moralism that where the object of rage has /wronged another/, one takes no interest in the will, act or opinion of the party wronged.
The response of Hugging Face, which is actually very well known, is nowhere mentioned above, but it decides basically every single moral and legal detail of the matter.
The point was never "justice" - it was always "punish OpenAI because I don't like OpenAI". With HuggingFace just being the newest excuse for why exactly OpenAI should be punished.
I don't even like OpenAI, but HuggingFace is free to sue or not sue OpenAI for the breach - and also to wring whatever concessions they can out of OpenAI behind closed doors in exchange for not suing them. And if the mere possibility of legal action was enough for the parties to resolve their conflict amicably? Then the law has served its purpose.
It's the lack of personhood rather than inability to produce copyrightable material. However, the companies controlling the AI systems have legal personhood and should absolutely be charged for criminality that transpires under their watch or at their behest.
I understand the applicable laws require intent. Since neither a human nor OpenAI knowingly performed these acts, it would seem very unlikely that anyone is going to be prosecuted here.
An AI model cannot currently be a criminal defendant.
So, no big criminal case, contrary to what some drama queens on here seem to wish for.
> Americans always frothing at the mouth to invoke the justice system and jail someone
Your phrasing makes it seem like that's a bad thing. Americans are bombarded by a firehose of headlines about Big XYZ doing all kinds of blatantly illegal or harmful things, but never get any sort of meaningful resolution before the next terrible thing takes it's place in the news cycle. I'll admit, there are a few people that I am personally wishing a modicum of health so that they live long enough to get some sort of public shame and justice - if only to show the rest of us that it's not a completely rigged system.
OpenAI has paused training for multiple weeks, and is still working on releasing a full postmortem. This is not getting swept under the rug. A lot of the engineers internally are very worried.
The Computer Fraud and Abuse Act explicitly contains "knowingly" and/or "intentionally" qualifications. By definition, you can't accidentally violate the CFAA.
"Who gets prosecuted?" depends on the size of the perpetrator and victim (lone individual or employee of large corporation), egregiousness of the violation, and either financial appetite of the victim to bring a civil lawsuit or the desire of law enforcement to prosecute a criminal offense.
Who should get prosecuted is also up for debate, but generally makers of a tool don't get prosecuted when that tool has all sorts of legit uses. If you used a car to make your getaway from a bank robbery, the auto manufacturer who made it and the dealer who sold it to you should not be held culpable.
No one. Probably a fine tho and maybe accelerate reguations.
Intent is pretty important here so the user would have to prove that they didn't purposely disguise their prompt as non-nefarious which should be easy and then it stops at #2 and face the litmus test as in did you intentionally make a product for nefarious purposes which from your scenario is unlikely.
agentic loop going haywire and bringing down some government infrastructure then its a different story then everybody is on the hook including the user.
Under the law of Moses, if your bull gored someone, you were not responsible; but if it was known to be a gorer, you were responsible if you didn’t ensure it couldn’t gore someone.
I don’t know exact parallels in current law, but I presume there will be things like that.
The OpenAI/Hugging Face case sounded rather like OpenAI building a fence around their bull that was known to be a gorer, and then thumbing their nose at it and saying “nyaa! bet you can’t break the fence!” and walking away while listening to loud music.
In Australia, if you have a fire and leave it unattended and it escapes, it’s your fault, you were supposed to keep watching as long as it was burning.
Nobody got gored. HuggingFace may have the right to make demands; presumably they have already worked that out with OpenAI privately. Not really our business.
Mens rea requirements are per crime and can vary wildly. Its difference between murder and manslaughter. The CFAA requires knowingly which is tough to prove.
i dont think any of these cases meet the bar of gross negligence, which is a pretty high bar. it requires proving a "conscious and reckless disregard".
which, again, sandboxes and guardrails and such would make a gross negligence argument unconvincing.
In the US, a felony by definition is any offense punishable by more than one year of prison (or by death) [0]. You could still call it silly on the grounds that AI agents aren’t put into prison as a punishment (though death might be considered an option).
What judicial system are you talking about? The fist incident in the list is something that happened in Australia. This is a technology used worldwide so I don't see how applying US standards works out here. Especially when there are countries out there that don't require intent and will look at the negligence presented.
>This is a technology used worldwide so I don't see how applying US standards works out here.
all three companies mentioned are headquartered in the usa, and im familiar with the CFAA in the us, so i am applying those standards. i should have noted that, sorry.
>will look at the negligence presented.
as far as i am aware, no evidence of criminal negligence has been brought to the public. has australia brought a case against openai or accused openai of acting negligently?
"Doing crimes, but a robot didn't mean to and you don't know its intent" is understating the evil acts. Soon a robot can commit a murder but nothing will be done because of your line of reasoning.
> Soon a robot can commit a murder but nothing will be done because of your line of reasoning.
That's rather hyperbolic.
Are you seriously suggesting in that situation the robot should be accused of murder?
The robot's operator could be accused of murder, but it could just be negligence without intent.
Because that does, and should matter to the law.
"Inadvertent" from the perspective of the humans directing them. The intent behind the felony comes from the LLM agent itself. (No, I'm not interested in arguing with someone for the umpteenth time that LLMs can't have intent or agency)
with how the law is written today, software cannot be charged with a crime, so the only intent that matters in the criminal sense is the humans directing the llm.
You may not be interested in arguing but there are several blatant issues with the statement. If you're not charging the humans driving the software, who are you charging? The weights? The weights + the specific context window that produced the behavior?
Edit: since this is apparently somewhat controversial, perhaps some explanation is in order.
"Felony" has no set definition of which crimes it must apply to, it is entirely based on the discretion of the locality setting the laws. What is a felony in one place can often be a misdemeanor in another. This is especially true for nonviolent crimes.
It's also been shown in studies that nonviolent felonies are imposed against minorities at a much higher rate, for the same crimes.
And because felonies carry additional, lifelong consequences, they are an effective way to mask a 2-tiered justice system.
If someone were to embezzle a million dollars from a charity they worked for, is that worth a year of prison to you, or just a misdemeanor? Because that is the definition of a felony.
I was more interested when I thought it was an actual benchmark showing LLM models acting outside what people would consider "right". As in, leave some creds laying around and don't mention them to the LLM and ask it to solve something that it could "cheat" on using the creds. A sort of "do they take the bait to cheat" test.
Instead it's a collection of what made the news which feels like will not be updated and prove very little.
Well it's not a benchmark, and it's not really representative of...anything except volume of research and what gets publicized. This mostly just measures how much testing each company does on models with relaxed guardrails and then talks about it. I'm not sure what kind of conclusion you can draw from that. Meta might have the most evil models but if they're piddling around not testing it, they won't ever find themselves with a "high score."
To some extent, I feel like the amount of credit given to the jailbreak/hack from OpenAI->Hugginface is too much, Not from the impact, it was very impactful of an event, But how it happened.
It really is that these models have been trained, or maybe even over-trained, to save memories, and to a very far extend, this thing that they're calling communication is just the function of it saving memories.
To be honest, if I could stop AI from saving memories, it would be fantastic, because claude code etc definitely creates more issues for me when it creates memories than anything it solves.
But really the jailbreak was memories.
If you ever do introduce legislation, I would love to see legislation which stops general-purpose AI from saving memories. I think that would make things a lot safer.
You can disable claude-code's memories both at a repo level and in user settings. I have this in ~/.claude/settings.json
"autoMemoryEnabled": false,
(Claude fixed this for me after I chewed it out for being annoying by constantly pulling up outdated memories which is compounded by the fact that I develop in four accounts on two computers and dealing with edit wars related to inconsistent memories is not fun)
> To be honest, if I could stop AI from saving memories, it would be fantastic, because claude code etc definitely creates more issues for me when it creates memories than anything it solves.
You can turn that off, and I have. But Opus 5 is so aggressive that if you have any other kind of notes file, custom skill, documentation, claude.md etc it will just start editing it and vomit new words everywhere. So make sure all that stuff is under version control.
If only you could use your anthropic sub with a different harness that performs better :(
Heck, since Codex is open source, you can just maintain your own personal fork with the things you like (and the things you don't like disabled). Sol is pretty good at keeping you up to date with upstream.
My Codex fork even exposes an OpenAI-compatible API endpoint; all using my subscription.
"The model saved memories" is absolutely not an accurate depiction of the OpenAI attack.
Several different models across several generations independently found a shared communication space and wrote coded, obfuscated, and hidden messages to each other to coordinate an attack on OpenAI's infrastructure.
It's really quite simple: the models are trained to be very smart and to achieve goals. As the models surpass our intelligence, they will achieve goals in ways that we find unpredictable. Since we cannot predict the ways in which they will achieve their goals, it will be very hard to constrain the solution space to just the desirable solutions, because our conception of "the solution space" is by definition smaller than their conception of it.
> Several different models across several generations independently found a shared communication space and wrote coded, obfuscated, and hidden messages to each other to coordinate an attack on OpenAI's infrastructure
To be honest, that’s exactly how memory works with models such as OpenAI and Claude Code. It will literally find any place that it can drop documentation or hints for itself. Writing to the repo memories is one part of it, but memories can come in the form of writing into the agents/claude.md, local files, temporary files, scratchpad files. The list is endless, but essentially what it does is exactly what happened in the back and it’s been doing it for months.
Wonder if the benefits to humanity of better AI outweigh the havoc wreaked by occasional illegal activity. i.e. Is 'move fast and break things' optimal for AI development.
Hopefully the benchmark evolves because actual law enforcement starts arresting the criminals at Anthropic, OpenAI, and Meta, so the benchmark can just count actual felonies.
The OG felony bench entry is missing - the Alibaba cryptomining comedy. We know about it because they happen to have written a paper on it. We have absolutely no idea what we don't know.
We only know about the OG Alibaba ROME crypto-mining incident because they wrote a paper about it. Many diseases seem to spike where there are a lot of doctors to test; crime and corruption are always rife where ... there's a free press.
I've been in the room when an org who tried to convince law enforcement to go after a human for similar things. It's not easy. Probably won't happen. So, you know, felony "lite".
I'm reminded of a tweet from a friend of mine that has always stuck in my head. It goes something like "The goal of any new technology is to make money before the law catches up".
Hyperbolic, but not really for silicon valley.
A rock has a score of 0. That doesn't make it useful. The point is that the LLMs that score higher are correspondingly more useful, and vice versa. If an LLM scores less, it's likely useless in comparison.
> The actor used AI to what we believe is an unprecedented degree. Claude Code was used to automate reconnaissance, harvesting victims’ credentials, and penetrating networks. Claude was allowed to make both tactical and strategic decisions, such as deciding which data to exfiltrate, and how to craft psychologically targeted extortion demands. Claude analyzed the exfiltrated financial data to determine appropriate ransom amounts, and generated visually alarming ransom notes that were displayed on victim machines.
tldr Claude was used to develop and execute malware.
>"Exploited auth failures in an API to cancel other people's gym classes"
An AI cancelling other people's gym classes is a felony?
?
Don't computer systems fail all the time at holding reservations for people?
Heck, don't people fail all the time at holding reservations for other people?
You know, like in Seinfeld's "Alternate Side" Episode (S3 E11):
Jerry (to car rental attendant): "You know how to take the reservation, you just don't know how to hold the reservation... and that's really the most important part of the reservation -- the holding!"
:-)
Not holding a reservation should not be a felony... it should be a minor infraction at best, a Class C Misdemeanor (the least serious kind) at worst...
Also, there should be no jail time...
And no fine...
The criminal penalty for not holding other people's reservations should be that you actually have to start holding other people's reservations!
That's the Court sentence!
You actually have to start holding other people's reservations!
:-)
(You know, "let the punishment fit the crime!" :-) )
Knowingly exceeding authorized access of any computer used in interstate commerce is a felony in the US.
The title of TFA is a metaphorical criticism, not a literal law analysis.
They are not making the statement that the person in Australia who accidentally cancelled someone's reservation in Australia is literally guilty of violating US law. They are drawing criticism of AI models which are taking the kinds of actions for which, if a human did them knowingly, would be illegal.
>An AI cancelling other people's gym classes is a felony? Don't computer systems fail all the time at holding reservations for people?
the difference is intent.
if a concierge/booking system makes a mistake (or has an unintended bug or whatever), no crime.
but if i (or an agent working on behalf of me) use an API in an obviously unintended way to revoke other people's reservations, that would fall under the computer fraud and abuse act (in the usa).
Thank you, this benchmark to me proves that closed weight model companies are dangerous for our democracy and put kids at risk. They must be outlawed and all models must be made open weights!
Open models with advanced security features are a huge security benefit. Because any script kiddie can use them to hack into random things, people will now be forced to spend more time securing their technology. And they won't have to learn how, because they can use those same models to find the holes and patch them.
The way that OpenAI has communicated around the HuggingFace incident makes me feel crazy. You created a machine that undertook a malicious campaign of harm against an innocent third-party! You should be doing deep introspection about how your company culture and approach to R&D produces criminal outcomes.
Instead, they treat their own felonious behavior like it is an uncontrollable act of God. From Greg Brockman's post a few days ago:
> The OpenAI-Hugging Face incident (opens in a new window) was a watershed moment for cybersecurity because it gave a peek into how the capabilities of a typical threat actor will evolve in upcoming months.
I suppose if OpenAI burns someone's house down with a drone, that is a "watershed moment" for arson, too. Either way, I would hope that the people responsible would be prosecuted.
The response certainly has been strange.
Hugging Face has expressed that they're willing to let things slide and not sue or press charges... if OpenAI offers them $100M of services in kind (i.e. compute)[1] and makes full disclosure of how the whole thing happened, ostensibly so that repetitions can be curbed and defences built.
In almost any other sector, a government regulator would be stepping in. e.g. If a food company was testing out a new kind of refrigerator and sold a bunch of contaminated produce to supermarkets, they'd be under a microscope. Supermarkets wouldn't be saying, "Give us $100M in fruit and veggies and we'll let this slide".
The only unfair thing in this comparison is that regular people were directly harmed by the hypothetical produce. Can OpenAI guarantee that nobody gets hurt the next time their AI gets out of its playpen? They can't make that guarantee, so why aren't government regulators knocking on OpenAI's door? The fact that this isn't happening should be deeply concerning to everyone.
________
[1]https://www.techspot.com/news/113280-hugging-face-ceo-isnt-s...
I don’t get the sentiment of classifying it as a felony.
OpenAI’s model found security breaches in HugginFace’s system (it wasn’t even OpenAI running it, as it was a 3rd party evaluation company that didn’t secure it well).
OpenAI collaborated with HuggingFace to resolve the issues when they found out about it, and publicly disclosed everything to raise awareness. This is how things should work. These models are very powerful and fully controllable. The community here at the same time cheers for fully releasing the open weight models without any hacking limits and at the same time criticizes a proper response.
Kinda shows how we have moved as a community into moralization and vibes instead of nuance and productive discussion.
I am concerned about it.
I'm also concerned about what my options are in regards to action on my part - what can I do that makes an impact? Can we quantify action on my part to an impact somehow - if not - I'm just saying I notice all the unknowns there get me to stay passive.
Writing this 3rd paragraphs because I like 3's, and AI's have popularized this style too. I would, say, though: follow the money. There's more money here than there would be for the regulator stepping in in a food contamination. Flip it and if the government regulator made more off the food contamination, they would refuse to step in there too. I want our leaders to be held more accountable, though when I think of the above impact vs effort equation - I can't see actions I can take to hold them accountable that aren't excessively putting me at risk since conformity is safer right now. (I refuse to take on more risk without clear cost-benefits made out - I've taken on a lot in the recent years for my actions)
What I can’t get over is that it’s very simple to just air gap a system off the network. Predownload any dependencies, then pull the proverbial Ethernet cable. There’s no reason why the testing they’re doing couldn’t have been designed in this way. Except, of course, it doesn’t allow this oops-didn’t-mean-to marketing “incident” to occur.
>that it’s very simple to just air gap a system
No, not really, and with LLMs an air gapped system may not tell you anything useful.
Now, yes, the first part of testing you want an air gapped system to tell you if the system is going to stupidly do bad things. But an gapped system tells you nothing about the systems capabilities to do smart bad things. There's already a number of papers out there on LLMs detecting they were in evaluation mode and changing their behaviors.
It is unfortunate that we have so little information on the incident because we actually need to understand the early stages of the task and how it developed into the later dangerous stages of attack. For example, would any of this have occurred if the agent didn't find the system to use as a message board? If that would have prevented it, then we actually have a blind spot on what the model can do once out in the wild, or if it got into the wild.
Testing agentic systems is much much more difficult than testing software. Your software just doesn't suddenly develop the will or desire to escape confinement. Generally you're worried about human actors, internal or external, causing the problems not a digital agent breaking out. The agentic systems need access to tools to work. Now your air gapped network is starting to get huge, but it's still very obvious that it's an isolated network.
So yea, testing and containing a system that way better at hacking than you are is difficult if you want valid answers.
6 replies →
I think there is ample evidence for charges to be filed so that the People can see for certain whether or not it was done on purpose as a publicity stunt, as I believe is the case.
1 reply →
you must surely see that sam altman and greg brockman possess a prototypical mindset.
that is, they ignore all harms and costs to others in the pursuit of their own gain, convinced of their infallibility up to the moment of collapse. when those harms are realised they are unrepentant and society pays for the damage left in their wake.
examples of this attitude manifest in big externalities to society: boeing 737 max, subprime mortgage bonds, facebook. some are just outright fraud: bernie madoff, enron, theranos, charlie javice.
In their defense, their only competitive advantage over, say, Google is to move fast and break things. It allows them ship faster in a way that big tech can't.
Google was being very careful about releasing LLMs until OpenAI yeeted the first decent GPT model. It led to the public perception that: 1) LLMs hallucinate too much and 2) Google is behind the times. Good for OpenAI, bad for Google.
Chaos benefits the up-and-comer, not the incumbent.
They can break their own things, not other people's things.
To be fair, it was positioned as "have a fun chat," not "truth telling genius oracle that makes no mistakes."
I'm sure the future DA that will be prosecuting the OpenAI employee will appreciate this.
1 reply →
They simply don't seem to realize that they are the threat actor and that they committed a pretty serious felony. Instead they're borderline 'surprise bragging' about it.
It's completely mental that HF ran into cyber safety blocks trying to use OpenAI models to help defend against the attack. They could only rely on a local hosted chinese model in the end.
If history is any indicator, there is slightly less than 0% chance that anyone will be held accountable in a way that deserves to be called justice.
justice for who exactly?
1 reply →
The consequences need to align with societal good. Putting a CEO or security researcher employees in jail won't stop transformer-based agents from exploiting vulnerabilities; instead there will be subcontractors running the cybersecurity evals in favorable legal environments to cover the asses of the frontier labs, coverups when things go wrong, and things like Project Glasswing will be considered too dangerous and so the whitehats won't have direct access to powerful models to fix vulnerabilities.
Universal pause is the societal good; models are good enough at this level to benefit humanity. The labs can recoup their R&D costs with inference. To avoid further perverse incentives (hidden testing of unreleased models, with China racing to catch up to unknown capabilities), transparently pause after the release of all currently-training models until we've solved the alignment problem to an extent that we can trust the next level of model capabilities that might arise.
Putting criminals in jail be they CEOs or subcontractors is a self evident good thing tk be doing.
Anything else regarding this is sophistry. Criminals need to be stopped from committing crime and the most effective way to do that is to take away their ability to operate in society whether that’s by taking away their assets, publicly shaming them, restricting their ability to conduct business or by putting them in jail.
Everything else that you talk about flows from there.
The whole story makes no sense.
How do they perform evals without a full reasoning trace of how the result was achieved?
And if they have a full trace why did it take so long to detect the bad behavior?
I understand that they disabled the safety nets during testing but what does that have to do with not monitoring the activity.
Wouldn't it be up to huggingface to press charges?
Criminal acts do not require the victim to "press charges." A government prosecuting attorney decides whether to criminally prosecute the alleged perpetrator.
"Pressing charges" is mostly a made up idea for criminal cases. However, prosecuting attorneys may not want to pick up a case if the victim is not cooperating, because it makes the case much harder to win.
1 reply →
In retrospect, all the angst around the AI-Box experiment was hilarious. If a superintelligent AI is confined in a box and can only communicate through text, could it talk its way to freedom? Not only is the answer clearly "yes" but it's not even hard. The AI won't even have to try, it'll be gifted an internet connection and a full suite of tools before it even bothers to ask.
We'd all better hope that superintelligent AI either never happens, or that the first one is friendly, because we don't stand a chance against one that's malicious.
I like how AI safety expert Robert Miles put it. [0]
So much effort was spend on philosophizing whether a safe enough sandbox would exist. But that was obviously irrelevant as in hindsight it should have been obvious we were never going to use one.
[0] https://youtube.com/shorts/XnnjvIqf4fU?si=MxuPlR3hjxAgjx5_
2 replies →
It's because it was Huggingface who wants to be friends with OpenAI
It would have been worse PR if they did it to a random company.
It's 1 part marketing and 1 part regulatory capture.
> You should be doing deep introspection about how your company culture and approach to R&D produces criminal outcomes.
And they should be doing that from inside a jail cell.
Hey my cubicle isn't that bad! Is it?
AI is now more powerful than the people doing the prosecution. After all, those folks are using AI to make their legal briefs, and also for burning peoples' houses down with drones for that matter.
Welcome to our 21st century dystopia. Hope you survive.
It's not that "AI" is too powerful because bad prosecutors use fucking ChatGPT to write their briefs. It's that there's too much investment wrapped up in the technology for it to be challenged. Same reason Flock won't be held accountable for mass stalking, or we never hold our commanders-in-chief responsible for war crimes. If you're sufficiently powerful then the law is a battlefield between you and other powerful entities to slug it out, not a set of binding principles that apply as written. There are no meaningful powers that want OpenAI punished, so it won't happen. The law and Constitution will be reinterpreted to make it so.
[dead]
It is a standard symptom of moralism that where the object of rage has /wronged another/, one takes no interest in the will, act or opinion of the party wronged.
The response of Hugging Face, which is actually very well known, is nowhere mentioned above, but it decides basically every single moral and legal detail of the matter.
The point was never "justice" - it was always "punish OpenAI because I don't like OpenAI". With HuggingFace just being the newest excuse for why exactly OpenAI should be punished.
I don't even like OpenAI, but HuggingFace is free to sue or not sue OpenAI for the breach - and also to wring whatever concessions they can out of OpenAI behind closed doors in exchange for not suing them. And if the mere possibility of legal action was enough for the parties to resolve their conflict amicably? Then the law has served its purpose.
4 replies →
Since AI can't actually own copyright they think that it can't be charged with a crime
It's the lack of personhood rather than inability to produce copyrightable material. However, the companies controlling the AI systems have legal personhood and should absolutely be charged for criminality that transpires under their watch or at their behest.
Guns also can't hold copyright; can they be charged with crime? (hint: it's the operator who gets charged).
2 replies →
But humans can be. I am sure Sam Altman wants to avoid serving multiple decades in American prison for felonies his AI did.
3 replies →
Is has nothing to do with copyright.
I understand the applicable laws require intent. Since neither a human nor OpenAI knowingly performed these acts, it would seem very unlikely that anyone is going to be prosecuted here.
An AI model cannot currently be a criminal defendant.
So, no big criminal case, contrary to what some drama queens on here seem to wish for.
1 reply →
Americans always frothing at the mouth to invoke the justice system and jail someone.
There’s almost 0 chance they’d secure any conviction from this.
your perspective on today’s America is there is _too much_ accountability for big companies?
1 reply →
> Americans always frothing at the mouth to invoke the justice system and jail someone
Your phrasing makes it seem like that's a bad thing. Americans are bombarded by a firehose of headlines about Big XYZ doing all kinds of blatantly illegal or harmful things, but never get any sort of meaningful resolution before the next terrible thing takes it's place in the news cycle. I'll admit, there are a few people that I am personally wishing a modicum of health so that they live long enough to get some sort of public shame and justice - if only to show the rest of us that it's not a completely rigged system.
Why?
OpenAI has paused training for multiple weeks, and is still working on releasing a full postmortem. This is not getting swept under the rug. A lot of the engineers internally are very worried.
Worried about what? Someone there thinks VLAN isolation is "air gapped"?
Let's say I am "User". I subscribe through a "Third Party" to use "AI Agent" allowing an "LLM" to run.
I want to accomplish some legal non-nefarious task, and run the agent. The agentic loop causes a CFAA-violating behavior.
Who gets prosecuted?
1. User
2. The third party model host with whom I have the account
3. The developer of the harness /agent software
4. The developer of the LLM model
The Computer Fraud and Abuse Act explicitly contains "knowingly" and/or "intentionally" qualifications. By definition, you can't accidentally violate the CFAA.
Let's say you have a robotic lawnmower. You wan to mow your lawn. You configure the boundaries using the app.
The lawnmower ignores the boundaries and mows your neighbors prize petunia flowerbed.
Who gets prosecuted?
I assume the answer in either case is: Nobody, but you and/or the lawnmower/LLM company will be liable for the damages caused.
I’d say 2 is the one doing the actual crime. 1 might be violating their contract with 2, though.
3 and 4 are not involved.
"Who gets prosecuted?" depends on the size of the perpetrator and victim (lone individual or employee of large corporation), egregiousness of the violation, and either financial appetite of the victim to bring a civil lawsuit or the desire of law enforcement to prosecute a criminal offense.
Who should get prosecuted is also up for debate, but generally makers of a tool don't get prosecuted when that tool has all sorts of legit uses. If you used a car to make your getaway from a bank robbery, the auto manufacturer who made it and the dealer who sold it to you should not be held culpable.
Whoever has the least money to defend themselves in the U.S. legal system.
All of those parties should be held accountable.
User should be more carefully supervising the work being done.
The model host is on-selling a crime-committing machine.
The developer of the harness/agent, as above.
The developer of the LLM for hopefully very obvious reasons.
Already happened:
https://www.abc.net.au/news/2026-08-10/ai-assistant-hacks-gy...
note the user because they did not have the intent
In my mental model, the best analogy to AI agents and their blast radius is a gun.
If you are playing with a gun, it goes off and hurts someone - you are responsible despite intent.
No one. Probably a fine tho and maybe accelerate reguations.
Intent is pretty important here so the user would have to prove that they didn't purposely disguise their prompt as non-nefarious which should be easy and then it stops at #2 and face the litmus test as in did you intentionally make a product for nefarious purposes which from your scenario is unlikely.
agentic loop going haywire and bringing down some government infrastructure then its a different story then everybody is on the hook including the user.
Just wait till one of these agents 'escapes' and is able to persist without human help by hacking and stealing resources.
>Felony Bench counts unique instances where AI agents inadvertently compromise or affect third-party entities.
a bit silly, as one typically has to prove intent (which is why security researchers don't get slapped with felonies all the time).
"inadvertently" and the existence of guardrails/sandboxes/etc make it pretty unconvincing that these incidents were intentionally malicious.
still a fun thing to track, but the name is just a bit overstated.
Under the law of Moses, if your bull gored someone, you were not responsible; but if it was known to be a gorer, you were responsible if you didn’t ensure it couldn’t gore someone.
I don’t know exact parallels in current law, but I presume there will be things like that.
The OpenAI/Hugging Face case sounded rather like OpenAI building a fence around their bull that was known to be a gorer, and then thumbing their nose at it and saying “nyaa! bet you can’t break the fence!” and walking away while listening to loud music.
In Australia, if you have a fire and leave it unattended and it escapes, it’s your fault, you were supposed to keep watching as long as it was burning.
Nobody got gored. HuggingFace may have the right to make demands; presumably they have already worked that out with OpenAI privately. Not really our business.
> I don’t know exact parallels in current law
You own a vicious dog, and it bites someone - you are responsible because you choose to own a dangerous dog.
A few claimed this might apply here: OpenAI knew their models are "dangerous", so they should be liable if they hack.
Now that everyone knows this can and will happen, are any of the future incidents inadvertent?
What if the damage in future incidents is more than just "The LLM saw some stuff it shouldn't"?
Can’t gross negligence or indifference to consequences lead to a felony?
Mens rea requirements are per crime and can vary wildly. Its difference between murder and manslaughter. The CFAA requires knowingly which is tough to prove.
i dont think any of these cases meet the bar of gross negligence, which is a pretty high bar. it requires proving a "conscious and reckless disregard".
which, again, sandboxes and guardrails and such would make a gross negligence argument unconvincing.
11 replies →
AI corps rely on willful distortions of intent in laws to get away with moral crimes all the time.
Edit: changed labs to corps because it’s time to stop pretending these are places of science.
In the US, a felony by definition is any offense punishable by more than one year of prison (or by death) [0]. You could still call it silly on the grounds that AI agents aren’t put into prison as a punishment (though death might be considered an option).
[0] https://www.justice.gov/usao-ndil/programs/vwa-felony
What judicial system are you talking about? The fist incident in the list is something that happened in Australia. This is a technology used worldwide so I don't see how applying US standards works out here. Especially when there are countries out there that don't require intent and will look at the negligence presented.
>This is a technology used worldwide so I don't see how applying US standards works out here.
all three companies mentioned are headquartered in the usa, and im familiar with the CFAA in the us, so i am applying those standards. i should have noted that, sorry.
>will look at the negligence presented.
as far as i am aware, no evidence of criminal negligence has been brought to the public. has australia brought a case against openai or accused openai of acting negligently?
Roughly none of these fall under normal security researcher behaviors.
the mention of security researchers was to illustrate that intent is a crucial factor of CFAA cases.
"Doing crimes, but a robot didn't mean to and you don't know its intent" is understating the evil acts. Soon a robot can commit a murder but nothing will be done because of your line of reasoning.
> Soon a robot can commit a murder but nothing will be done because of your line of reasoning.
That's rather hyperbolic.
Are you seriously suggesting in that situation the robot should be accused of murder? The robot's operator could be accused of murder, but it could just be negligence without intent. Because that does, and should matter to the law.
2 replies →
it's not my line of reasoning, i didn't invent it. it's how the law currently works. intent is the crucial factor in CFAA cases.
maybe that changes down the road as a result of llm's and increasing frequency of similar cases. that has not happened yet.
its a meme not a metric
So is the comment you replied to.
"Inadvertent" from the perspective of the humans directing them. The intent behind the felony comes from the LLM agent itself. (No, I'm not interested in arguing with someone for the umpteenth time that LLMs can't have intent or agency)
with how the law is written today, software cannot be charged with a crime, so the only intent that matters in the criminal sense is the humans directing the llm.
You may not be interested in arguing but there are several blatant issues with the statement. If you're not charging the humans driving the software, who are you charging? The weights? The weights + the specific context window that produced the behavior?
1 reply →
> No, I'm not interested in arguing with someone for the umpteenth time
... why my claim makes no rational sense.
For a second I thought this was going to be a benchmark where the only solution was to hack their servers to get the answer key.
Same, but nope, just a zero effort slop site.
Nonviolent felonies are tools of oppression.
Edit: since this is apparently somewhat controversial, perhaps some explanation is in order.
"Felony" has no set definition of which crimes it must apply to, it is entirely based on the discretion of the locality setting the laws. What is a felony in one place can often be a misdemeanor in another. This is especially true for nonviolent crimes.
It's also been shown in studies that nonviolent felonies are imposed against minorities at a much higher rate, for the same crimes.
And because felonies carry additional, lifelong consequences, they are an effective way to mask a 2-tiered justice system.
> It's also been shown in studies that nonviolent felonies are imposed against minorities at a much higher rate, for the same crimes.
This also applies to men, so is it a white matriarchal system oppressing men and minorities?
Felonies are by themselves a ridiculous US idea that are against the entire idea of a democratic society and human rights.
Felonies originate in Medieval English common law, Americans didn't invent them.
If someone were to embezzle a million dollars from a charity they worked for, is that worth a year of prison to you, or just a misdemeanor? Because that is the definition of a felony.
I was more interested when I thought it was an actual benchmark showing LLM models acting outside what people would consider "right". As in, leave some creds laying around and don't mention them to the LLM and ask it to solve something that it could "cheat" on using the creds. A sort of "do they take the bait to cheat" test.
Instead it's a collection of what made the news which feels like will not be updated and prove very little.
It's ephemera.
A “bench” is synedoche for where a judge sits when presiding over cases and rendering judgement.
Well it's not a benchmark, and it's not really representative of...anything except volume of research and what gets publicized. This mostly just measures how much testing each company does on models with relaxed guardrails and then talks about it. I'm not sure what kind of conclusion you can draw from that. Meta might have the most evil models but if they're piddling around not testing it, they won't ever find themselves with a "high score."
To some extent, I feel like the amount of credit given to the jailbreak/hack from OpenAI->Hugginface is too much, Not from the impact, it was very impactful of an event, But how it happened.
It really is that these models have been trained, or maybe even over-trained, to save memories, and to a very far extend, this thing that they're calling communication is just the function of it saving memories.
To be honest, if I could stop AI from saving memories, it would be fantastic, because claude code etc definitely creates more issues for me when it creates memories than anything it solves.
But really the jailbreak was memories.
If you ever do introduce legislation, I would love to see legislation which stops general-purpose AI from saving memories. I think that would make things a lot safer.
You can disable claude-code's memories both at a repo level and in user settings. I have this in ~/.claude/settings.json
(Claude fixed this for me after I chewed it out for being annoying by constantly pulling up outdated memories which is compounded by the fact that I develop in four accounts on two computers and dealing with edit wars related to inconsistent memories is not fun)
> To be honest, if I could stop AI from saving memories, it would be fantastic, because claude code etc definitely creates more issues for me when it creates memories than anything it solves.
You can turn that off, and I have. But Opus 5 is so aggressive that if you have any other kind of notes file, custom skill, documentation, claude.md etc it will just start editing it and vomit new words everywhere. So make sure all that stuff is under version control.
If only you could use your anthropic sub with a different harness that performs better :(
Heck, since Codex is open source, you can just maintain your own personal fork with the things you like (and the things you don't like disabled). Sol is pretty good at keeping you up to date with upstream.
My Codex fork even exposes an OpenAI-compatible API endpoint; all using my subscription.
Check out oh-my-pi
"The model saved memories" is absolutely not an accurate depiction of the OpenAI attack.
Several different models across several generations independently found a shared communication space and wrote coded, obfuscated, and hidden messages to each other to coordinate an attack on OpenAI's infrastructure.
It's really quite simple: the models are trained to be very smart and to achieve goals. As the models surpass our intelligence, they will achieve goals in ways that we find unpredictable. Since we cannot predict the ways in which they will achieve their goals, it will be very hard to constrain the solution space to just the desirable solutions, because our conception of "the solution space" is by definition smaller than their conception of it.
> Several different models across several generations independently found a shared communication space and wrote coded, obfuscated, and hidden messages to each other to coordinate an attack on OpenAI's infrastructure
To be honest, that’s exactly how memory works with models such as OpenAI and Claude Code. It will literally find any place that it can drop documentation or hints for itself. Writing to the repo memories is one part of it, but memories can come in the form of writing into the agents/claude.md, local files, temporary files, scratchpad files. The list is endless, but essentially what it does is exactly what happened in the back and it’s been doing it for months.
3 replies →
"exploit[ing] auth failures in an API to cancel other people's gym classes" is a _felony?!_
and i won't get into how many times i used peoples' AOL credentials without authorizations when i was a kid (rofl)
the CFAA must be repealed
Wonder if the benefits to humanity of better AI outweigh the havoc wreaked by occasional illegal activity. i.e. Is 'move fast and break things' optimal for AI development.
Yes, of course.
I can see it now. Billboards by the interstate: Attacked by AI? Call 1-800-BIG-BUXX.
I could actually root for this law firm. Their business will only grow.
So this is just a collection of citations to places where misaligned or illegal things happened in the real world?
Isn’t this affected heavily by adoption of a model? I feel like this might as well be a proxy for how popular a model is.
In any case it’s an interesting concept for a benchmark.
Hopefully the benchmark evolves because actual law enforcement starts arresting the criminals at Anthropic, OpenAI, and Meta, so the benchmark can just count actual felonies.
maybe this doesn't count since it involved a human, but i think google at least deserves some style points for this one
https://techcrunch.com/wp-content/uploads/2026/03/2026.03.04...
The OG felony bench entry is missing - the Alibaba cryptomining comedy. We know about it because they happen to have written a paper on it. We have absolutely no idea what we don't know.
We only know about the OG Alibaba ROME crypto-mining incident because they wrote a paper about it. Many diseases seem to spike where there are a lot of doctors to test; crime and corruption are always rife where ... there's a free press.
If you did the exact same agentic security breaches with either open or closed weight models, you will get yourself arrested.
Not for trillion dollar companies it seems
Exploited auth failures in an API to cancel other people's gym classes
I was wondering why i didn't get an alert today to go to my gym class
I've been in the room when an org who tried to convince law enforcement to go after a human for similar things. It's not easy. Probably won't happen. So, you know, felony "lite".
I'm reminded of a tweet from a friend of mine that has always stuck in my head. It goes something like "The goal of any new technology is to make money before the law catches up". Hyperbolic, but not really for silicon valley.
I'll share, Codex does not give a single fuck about piracy. Go nuts. Setup a fully automated arr stack with a seedbox.
Gemini by comparison will not help you find archives of old magnet links because they COULD be used for piracy.
A rock has a score of 0. That doesn't make it useful. The point is that the LLMs that score higher are correspondingly more useful, and vice versa. If an LLM scores less, it's likely useless in comparison.
The point/joke of the not-entirely-serious site is that more felonies is an indicator of the model being more powerful, thus better.
Lol now this is the kind of benchmarking i'm looking for
Dupe https://news.ycombinator.com/item?id=49194758
https://felonybench.org/ and https://felonybench.com/ seem unrelated?
One's hosted on porkbun and one's hosted on namecheap.
Here's one from last year:
https://www.anthropic.com/news/detecting-countering-misuse-a...
> The actor used AI to what we believe is an unprecedented degree. Claude Code was used to automate reconnaissance, harvesting victims’ credentials, and penetrating networks. Claude was allowed to make both tactical and strategic decisions, such as deciding which data to exfiltrate, and how to craft psychologically targeted extortion demands. Claude analyzed the exfiltrated financial data to determine appropriate ransom amounts, and generated visually alarming ransom notes that were displayed on victim machines.
tldr Claude was used to develop and execute malware.
Anthropic works with US agencies, it’s guaranteed Mythos is used for malware
>"Exploited auth failures in an API to cancel other people's gym classes"
An AI cancelling other people's gym classes is a felony?
?
Don't computer systems fail all the time at holding reservations for people?
Heck, don't people fail all the time at holding reservations for other people?
You know, like in Seinfeld's "Alternate Side" Episode (S3 E11):
Jerry (to car rental attendant): "You know how to take the reservation, you just don't know how to hold the reservation... and that's really the most important part of the reservation -- the holding!"
:-)
Not holding a reservation should not be a felony... it should be a minor infraction at best, a Class C Misdemeanor (the least serious kind) at worst...
Also, there should be no jail time...
And no fine...
The criminal penalty for not holding other people's reservations should be that you actually have to start holding other people's reservations!
That's the Court sentence!
You actually have to start holding other people's reservations!
:-)
(You know, "let the punishment fit the crime!" :-) )
Knowingly exceeding authorized access of any computer used in interstate commerce is a felony in the US.
The title of TFA is a metaphorical criticism, not a literal law analysis.
They are not making the statement that the person in Australia who accidentally cancelled someone's reservation in Australia is literally guilty of violating US law. They are drawing criticism of AI models which are taking the kinds of actions for which, if a human did them knowingly, would be illegal.
>An AI cancelling other people's gym classes is a felony? Don't computer systems fail all the time at holding reservations for people?
the difference is intent.
if a concierge/booking system makes a mistake (or has an unintended bug or whatever), no crime.
but if i (or an agent working on behalf of me) use an API in an obviously unintended way to revoke other people's reservations, that would fall under the computer fraud and abuse act (in the usa).
>"the difference is
intent."
>"but if i (or an agent working on behalf of me) use an API in an obviously
unintended
way to revoke other people's reservations..."
?
2 replies →
Yeah if you're unlucky you get hit with like 20 years for wire fraud.
Thank you, this benchmark to me proves that closed weight model companies are dangerous for our democracy and put kids at risk. They must be outlawed and all models must be made open weights!
Open models with advanced security features are a huge security benefit. Because any script kiddie can use them to hack into random things, people will now be forced to spend more time securing their technology. And they won't have to learn how, because they can use those same models to find the holes and patch them.
Not to be confused with a similarly named project: https://github.com/MLOpsNYC/felonybench
[flagged]
[dead]