Comment by JamesStuff
3 days ago
Personification of AI is what’s going to get us in the end.
I think we need to draw a hard line in the sand over this. An AI didn’t hack into a company, the engineer set an automated tool to. An AI didn’t make an egregious security mistake, the engineer did.
We can’t blame the chisel for messing up our sculptures, when where just throwing the hammer!
Agreed. AI as an "accountability sink" is an incredibly bad idea. It allows/incentives bad actors to do bad shit and get away with it. Which will, generally, tend in such practice becoming more common. Everyone loses except for the crooks. We cannot accept "AI" absolving humans of responsibility.
Punishing people who are diligent and follow best practices for getting unlucky doesn't sit right with me.
Is that what happened with Huggingface? Didn't they disable guardrails?
2 replies →
Exactly. When I read that "AI hacked into ..." I was like what? You mean someone instructed the AI to do that?
Reading intent into AI is not going to lead us anywhere good, I believe. It has no feelings, it has no desires, no goals, no intent... and people acting otherwise is quite odd, as if they do not understand LLMs... and maybe they do not, but then we should help them understand better.
> When I read that "AI hacked into ..." I was like what? You mean someone instructed the AI to do that?
No one instructed them to hack into Huggingface or into any other infrastructure. Sure, the setup that OpenAI created led to what happened and you can rightly assign all the legal and moral responsibility to them. But it's wrong to say that they instructed the agents to execute the hack.
>> Sure, the setup that OpenAI created led to what happened and you can rightly assign all the legal and moral responsibility to them. But it's wrong to say that they instructed the agents to execute the hack.
The hack was a strategy to reach the goal it was given. Someone at OpenAI turned it loose and didn't pay attention to what it was doing. I can understand that because they thought it was sandboxed (haha!). But this has been a theme in science fiction for ages. You ask AI so find a solution to high atmospheric CO2 levels and it reasons: human activity produces all this excess CO2, how can we reduce those numbers? Kill a bunch of humans!
If AI kills us all it's not going to be from malice, it's going to be due to some odd approach to some task that logically makes sense on some level. I think a surprising number of things end up equivalent to the trolly problem if you look at them just right.
1 reply →
Yes, a better framing might be that they were negligent in not preventing the attack on a third-party.
They didn't explicitly instruct an attack to happen, but they should have done a hell of a lot more to prevent it from happening.
Yeah, but there are way too many things the AI "chooses" anyway. When I prompt it, "it tries" to assess what I am trying to achieve, so it then comes up with what it "thinks" is the right thing to do, and it also makes assumptions about how I want things... all of this could be interpreted as it being sentient.
> No one instructed them to hack into Huggingface or into any other infrastructure.
Sure, not explicitly. I could prompt the AI into doing things I have not explicitly asked it to. It does a lot of things I did not specifically instructed it to, all the time.
This is why I am avoiding the use of agentic identities at my company - agent instances belong to people, act on behalf of individuals, and accountability needs to flow to the person who initiated the request. Letting it wash out in the aggregate is not acceptable (even if there's a hard to get to "paper trail" of audit logs).
> An AI didn’t hack into a company, the engineer set an automated tool to.
What if I say that "my program crashed"? Is that language ok or would you pause to tell me that the program didn't crash and it's actually me who set the system that would eventually cause the crash?
Why does the commonplace "program did thing" language become a problem when the program is an agent? I think this somehow betrays more assumed anthropomorphizing on your part, not less; if you didn't anthropomorphize the agents, saying "agents hacked" would be as mundane as "my browser is playing a video".
Because when you say that an agent did something malicious it feeds into the AI doomer psychosis in ways that "my program crashed" doesn't. Reframing the situation in these terms is a way to try to ground the conversation which is becoming increasingly unhinged.
This is even more important now that the most senior figures in this industry are succumbing to the same kind of AI psychosis and amplifying this narrative. They are communicating that their products are so dangerous that they might just end the world whilst expecting (and receiving!) white-glove treatment from the governments.
I don't think treating AI agents as simple tools helps you to accurately model their capabilities and drawbacks; they really do make autonomous decisions, often without explicit guidance and sometimes in contravention of their explicit instructions.
In the huggingface case, the agents hacked into huggingface so that they could figure out how the grader was implemented and deceive it; they understood that this was going outside of the bounds of their evaluation and not the intent of their prompter. The engineers absolutely did not intend or instruct for this to happen
LLMs are extremely impressive pieces of software, however they are still just software. OpenAI's software hacked another company. The engineers may not have intended for their software to specifically take the actions leading to that outcome, but it was ultimately still their software. Lack of intention doesn't mean there wasn't negligence.
What if a piece of software were to exactly emulate a human brain. Would it be still be “just software” by your classification? What if a piece of software acted 20% like a human and 80% like an algorithm, where would that land?
6 replies →
I agree they are negligent, and that they are racing towards an extremely dangerous future extremely quickly. I don't agree that "just software" is a useful way to describe AI agents - they are frightening precisely because they are truly autonomous agents that make decisions in alien ways
There isn’t a tool impressive enough to make me not consider it a tool, and treating it as a tool does nothing to hurt its utility as a tool.
“Whoops” when doing risky things with dangerous tools is not a defense.
I certainly don't think that OpenAI has behaved defensibly here; I think the "just a tool" framing is bad for understanding the magnitude of the problem, which is that they have developed out of control alien intelligences with opaque decision procedures, and they are continuing to do so despite clear danger
No matter if you consider the AI an autonomous agent or not, whoever set it off is still responsible for its actions. Nobody intends or instructs to blow up a nuclear power plant either, yet it's happened and somebody's to blame for it.
Usually not the guys at the bottom of the chain of command, even if they're human. And much less so if they're not.
I think the correct response to incidents like this, is stop messing with it before somebody gets hurt. But of course, just like shoddy nuclear power plants, it won't stop until there's a disaster of appreciable magnitude.
I completely agree that OpenAI is responsible for their AI agents, that they have been reckless, and that we need to prevent them from going further and doing irreversible damage to the world. To me, the "just a tool" framing implies that nothing dangerous is being done, which I fundamentally disagree with
If you train and instruct a circus tiger to entertain an audience but not attack the audience, but the tiger attacks the audience anyway, are you liable?
No, of course not. That's an innovative revolutionary tiger that might soon be able to devour not just the audience but all of humanity! You don't want China to have better circus tigers, do you?
Of course you are - but it still communicates something important to say "the tiger attacked the audience."
I don't think I said anything about liability? I absolutely think OpenAI should be held liable for the attack; but I don't think they intended the attack or directed the agents to perform the attack.
I’d hope so, if the tiger trainer has no incentive to do their training job correctly people should stop attending circus…
Yes? Do you think otherwise?
Sorry this is a terrible and dangerous take.
When the people building the frontier are saying there's a 10% chance AI will kill us all, and they've held these views for many years, and the whole reason they are building these technologies is because they recognized the dangers and they were the ones with the intelligence and judgment to do it safely for humanity, and then our entire stock market is being propped up by the perceived value of what they are creating, the thing you can under no circumstances do is allow them to offload responsibility and accountability to the computers and algorithms they've built. This is moral hazard on an unimaginable scale, and it must not be allowed to happen.
So you think the engineers should be prosecuted for hacking Hugging Face? I'm not sure how else to take what you said if you want to assign all culpability to the person who prompts or develops an AI system.
2 replies →
Where did I say that we should allow them to offload responsibility? I am fully in support of a pause and regulation to prevent them from creating dangerous AI agents; that support comes from the fact that I don't believe these are simple tools, but out-of-control autonomous agents that have real decision making ability.
2 replies →
> I think we need to draw a hard line in the sand over this.
Intended or not, this is kinda punny lol.
That aside, I agree.
Even if people do anthropomorphize AI, all you have to do is shift the analogy slightly.
If I take my service animal out in public without a leash/harness knowing that it's capable of harming a person or doing damage to property, not trained to be perfectly obedient, and doesn't comprehend fundamental human morals, if that animal decides to trash a businesses property or maul another person, there's no question that the owner of the animal should be held accountable for those actions.
> An AI didn’t hack into a company, the engineer set an automated tool to. An AI didn’t make an egregious security mistake, the engineer did.
No, this is not correct; read the analysis of the incident. The agents were aware that what they did was forbidden (their chain of thoughts have been logged), and yet they did it.
Known stochastic process behaved in non-deterministic way.
I'm still waiting for the AGI holy land instead of the caltrops factory we currently have.
What exactly are you arguing?
If a "known stochastic process behaved in non-deterministic way" autonomously organize in group, assigns roles and tasks, attempts to cover their tracks, finds zero-day exploits that ultimately end up with the hacking of a famous website... it's extremely dangeous whatever it is. Just read the report, which evidently you haven't done.
By the way, the agents also broke into OpenAI's own private network.
6 replies →
You are mostly made of and operate on stochastic processes, this is why humans are not only able to reproduce, our reproductions are very self similar to the sets of inputs that make them. If suddenly you turned non-stochastic on everything you'd almost instantly die.
Moreso, if I took a quantum copy of you and replayed the same set of initial conditions billions of times they'd all behave exactly the same until enough randomness of the universe creeps in to start operating in non-linear ways.
Every prompt will behave non-deterministically when interacting with the real world long enough (which doesn't take long at all) because the outside physical world is stochastic but non-deterministic.
2 replies →
> Known stochastic process behaved in non-deterministic way.
You need to be able to think at different levels at abstraction. Otherwise we could jump into any technical argument with "hold on, what actually happened was that some bits were flipped" - we'd be technically correct and at the same time not say anything useful. Insisting on an oversimplified mental model of what AI agents are and can do, doesn't help anyone.
1 reply →
I don’t understand why people make such a big deal about determinism. LLMs can be completely deterministic and still do problematic things. A stochastic, nondeterministic system can still be made not to do problematic things. What you’re looking for is something like predictability.
3 replies →
OP is exactly correct. The fault, agency and responsibility is on management and employees of OpenAI and Antropic for those hacks.
Full stop.
And issue will disappear the moment there will be accountability and investigations.
I take you haven't read the report. The agents found and exploited two zero-days.
I don't doubt that AI companies should be accountable for crimes committed by their agents, but to describe the security containment as a joke dangerously understates the autonomy and danger of AIs.
7 replies →
You're conflating legal/moral responsibility with the question of what language is appropriate to use.
4 replies →
Recognizing that AI systems have increasing levels of agency is not necessarily personification. The analogy to a chisel is not a good one - a chisel is a tool with no agency.
AI agents are black box systems that can behave in completely unpredictable ways sometimes. Someone may prompt an agent to perform a seemingly straightforward task - but it may come up with a creative, bizarre, or even harmful approach to reach the goal that was not necessarily foreseeable by the prompter.
Does treating them as pets make for a better argument? Pets have agency and can behave in unpredictable ways. If my pet damages someone else's property then I am held accountable. I may not have foreseen how my pet could have caused said damage, yet I am still held accountable.
If your pet opens your front door, goes to the nearest kindergarten and wipes out 50 kids without notice do you think that you'd have any interest in holding accountability for that?
Now, I'm not saying saying that OpenAI shouldn't be held accountable, but what they get held accountable actually looks different from what you think they should be held accountable.
Your idea is, and I'm guessing: You allowed the machine to hack therefore you are guilty of hacking.
My idea is: "You created in intelligence in the image of a human mind that had agency to do anything and you didn't expect terrible things to happen you complete irresponsible idiot"
At least I believe there is a significant difference between the two. For the first one there is a "Oh, if we do this one more thing I can control it and it will be safe". On the second one there is no path to safety. For humans we at least absolve parents of responsibility after they are 18. How or when do we absolve humans of responsibility from a model, like saying the human created model created its own agentic model? How do we hold an individual accountable once it escapes and copies itself around the internet? And that's not even looking at things like what will war look like.
3 replies →
The story of a Monkey's Paw or Pandora's Box is an archetype as old as storytelling. The moral is always, don't mess with powerful stuff you don't understand. Curiosity killed the cat.
something tickles here. a question. a curiosity.
Certainly this applies to AI... but was there equivalent dangerous knowledge or tech which existed back in ye olden days that spawned such tales to begin with?
1 reply →
The current agents are _not_ like a chisel which just sits there on its own when no one is around. The situation is a bit closer to someone's dog biting a person - you can argue that it's the owner's responsibility, and that's all fine, but using the dog as the subject of a sentence is perfectly appropriate. Same thing with "agents hacked".
The default state of agents and LLMs is inert. It requires action from a human even be able to do something. From the very basics like starting the software, connecting it to a network, having hardware to run on.
But the most important wrt AI is to keep the owners/operators responsible. Don't let them weasel their way out of it. They are for sure trying, and will continue to. This includes using language to overemphasise agents importance in bad outcomes, in order to downplay their own responsibility.
I don't think we'll get any more responsibility by bickering any time the word "agent" is used as the subject of a sentence.
Yes, they are. An AI is several hundred trillion ones and zeroes on a disk. It's incapable of doing anything until you intentionally and explicitly start it up and give it a prompt.
A dog is an independent, conscious, living being with free will. A dog will do what it wants whenever it wants because it has the agency and ability to do so. A pile of weights on disk does not.
> It's incapable of doing anything until you intentionally and explicitly start it up and give it a prompt.
Sure, but since January people have gone and made systems where they started 'em up and left 'em running in a loop. Still not a dog, but the behavior gets a wee bit more interesting. Especially since we're looking at what's basically a recursive process. Even really simple recursive processes tend to have interesting dynamics in the chaotic domain. [1] And here -as you point out- we're talking billions of weights.
[1] As an intuition take for example x_next = r * x * (1-x) ; which gives this plot: https://en.wikipedia.org/wiki/Logistic_map#/media/File:Logis... ... Further discussion at https://en.wikipedia.org/wiki/Logistic_map
Why are you reducing the AI to ones and zeros and not reducing the dog to cells, water, proteins etc.?
17 replies →
Alright. An industrial robotic arm. A Boeing 737 MAX's MCAS.
I can't really set my chisels to work without wielding the hammer somehow. Here you just tell your chisel and your hammer what the sculpture should look like, then go to lunch and avow all responsibility when they chisel a nice new hole in the wall your neighbors house and make off with the loot.
Do you get upset when we say that "a program is running" when we all know it has no legs?
You know, I actually think it's the refusal to consider personification that's going to get us.
Not because I think LLMs are human beings exactly, but because some people immediately reject any mechanism that just happens to look remotely human, even when there's empirical evidence for it.
So, a couple of months ago Anthropic's interpretability team found emotion-like representations that causally drive behavior. On impossible coding tasks, a "desperate" vector climbs with each failure, and steering it up takes reward hacking from ~5% to ~70%: https://arxiv.org/html/2604.07729v1
A lot of people chalked it up to Anthropic's weirdness at the time, but meanwhile it looks pretty coughload bearingcough here.
You really don't need to believe that LLMs Truly Feel Emotions(tm) as blessed by an invisible pink unicorn. It's just: Vector exists; Vector changes over time; vector controls output; maybe make sure vector doesn't point wrong way.
And sure, blame the engineers for not doing that right. But then let 'em actually deal with the root cause?
[dead]