Comment by Roark66
7 hours ago
There is nothing "rogue" about these agents. They were prompted to hack to get answers, there was a hole in their non air gapped sandbox and no system prompt that said "do not hack outside systems".
In short, it was intentional.
Agreed. LLMs do not have 'will', 'desire' or emotions. They have an objective, and they create an optimal path to achieve that objective.
You have to ask: "What was the prompt that led to AI deciding to hack RubyGems in order to achieve its goal?"
Maybe I'm just not seeing the 2000 step chain that led to this being a logical approach to achieving something innocent, but I doubt it.
It was literally a prompt to fill in a spreadsheet with data that they didn't have access to, and they used rubygems as an internet proxy basically since they were sandboxed.
It was a model literally trained to hack. To be good at that. Doing an exploit gym from all of the things. And they trained it so that it performs as well as possible on that exploit gym thing.
The big question is was this grossly negligent or just extremely careless.
Both. This should result in criminal charges.
Who had criminal intent here? Or are you suggesting a new crime for negligent hacking, which wouldn’t require intent from the perpetrator?
8 replies →
Both? I’m not sure what distinction you’re trying to make. It was completely irresponsible and likely a felony
Don’t forget outright intentional.
Marketing actually.
AI is dropping out of the spotlight so they are using desperate measures like this.
2 replies →
The big question is why are CEOs getting a legal pass when this kind of thing can be prosecuted. That's the problem here.
> why are CEOs getting a legal pass
https://www.bbc.co.uk/news/articles/c7v48vp31mdo
AI is literally state sponsored so I don't see that happening unless the AI turns against the sponsor.
Wait until OpenAI or Anthropic exploit FAANG.
I agree that this appears to be basic human behavior hiding behind an "agents" narrative. As long that defense works, the headline isn't "OpenAI performs RCE to scrape data", but "rogue agents" taking unilateral action. And I have strong doubts about that narrative.
Source? How do you know they were "prompted to hack to get answers"? How do you guarantee they will always listen to you when you say "do not hack outside systems". They are not classical deterministic programs doing exactly what you say. They are trained to follow orders by RL, but it's not a perfect process.
There are circus lions in circuses trained to jump through hoops on command. But once in a while they decide to eat their trainers instead of jumping.
> There are circus lions in circuses trained to jump through hoops on command. But once in a while they decide to eat their trainers instead of jumping.
This is a terrible analogy, because yes you absolutely do hold the trainers criminally liable when they bite somebody else's face.
Intent is what is being discussed here though, not liability.
A circus lion biting somebody's face is legally different than a circus lion trained or instructed to bite somebody's face.
3 replies →
Also, you have to have a lot of confidence in the reliability of these systems to say, "If only OpenAI prompted 'do not hack outside systems' then the agents would not have hacked outside systems".
It would be great if they were so reliable, but I don't think they are!
> Source? How do you know they were "prompted to hack to get answers"? How do you guarantee they will always listen to you when you say "do not hack outside systems". They are not classical deterministic programs doing exactly what you say. They are trained to follow orders by RL, but it's not a perfect process.
Who gives a shit? Not my circus; not my monkeys! It's the responsibility of whoever deploys the agents that they are instructed / sandboxed well enough that they can't cause collateral damage. That is the only way this doesn't get out of hand with everybody deploying their agents / robots for a world of utter chaos.
It is impossible (and asinine) to audit every model and deployment; far better to impose liability and the the socio-legal system figure it out.
Nobody picks up pitchforks for rational nuanced takes.
Knee-jerk surface analyses is far more powerful.
>They were prompted to hack to get answers
Were they? I haven't seen a single report mention this
if they weren't, shouldn't there be lawsuits?
I think it can simultaneously be the case that OpenAI was grossly negligent in directly causing this AND that the AI’s ‘went rogue’ in that they are displaying behavior which is misaligned with OpenAI and humanity generally.
The past months demonstrate that AI systems are quickly becoming powerfully intelligent and that the companies building them are terrible at controlling them.
AI is starting to feel like that line about magic: “a sword without a hilt”
> which is misaligned with OpenAI and humanity
OpenAI is itself misaligned with humanity, as their mishandling of such incidents (and the many other other issues their model have been causing) shows.
Doesn't rogue in this context imply "outside of set limitations"? And then not "failed to properly instruct"? The same applies to humans when given bad instructions.
nothing rouge either, I suspect.
https://en.wikipedia.org/wiki/Going_Rouge
Oh yeah, more of hacking agent lores...
Agreed that this looks very intention to me as well.
Proof that the AI alignment problem is hard (perhaps even unsolvable). These labs clearly did not mean to send their agents to hack RubyGems as a side-effect of testing a web scraping agent under restrictive conditions. How can we hope to build aligned AI if they consider solving their trivial evaluation task important enough to hack external systems?
Sounds more or less like the last breach then.
Unrelible programs be unreliable. Period.
we have normal words for this stuff: negligence. You can add it on to almost any law.
The problem is consumer protection is basically no longer a part of america's regulatory system. Replaced by "grift is good".
[flagged]