Comment by tripzilch
14 hours ago
> They were not deliberately told to "go wild". The hacking wasn't even part of their test, it was the agents' attempt to cover up that they'd cheated on an impossible test.
TBH the more I read of these reports, the less I believe this.
These agents just weren't behaving in any way I've seen normal/publicly available agents do.
Sure I've heard (from other people, not seen myself) that they sometimes try to get around file system permissions or use `bash` to write when their `write` tool is disabled, or such.
But this is definitely another level, entirely.
There is this vague sense of desperation coming from many of these logs and I am sure they must have been motivated by something else, too.
We didn't see their system prompt or main prompt, right? We've only seen reports from what happened after deciding to break out.
OAI claims this was triggered by the task being literally impossible. That also doesn't quite add up, unless the other tasks that were possible, simply weren't hard enough? Otherwise wouldn't agents already start hacking when faced with a really hard task, too? Cause they wouldn't be able to differentiate. At least some of them would have started to somewhat poke their sandbox a bit?
Also I would have expected to see a few tens of other (perhaps less severe) public incidents from random people setting their models to YOLO, accidentally hacking stuff, this incident has been loud and messy enough, that if it happened to a few other people, we'd have heard about it.
Unless OAI's story is that it was specifically this batch of agents that crossed some threshold of going wild? (which would also raise some serious questions about how serious they take that danger ..).
Or maybe it is only dangerous if you have the compute resources to run 700 agents for weeks?
> There is this vague sense of desperation coming from many of these logs and I am sure they must have been motivated by something else, too.
If I had to guess, their motivation is "get reward for completing task". There's certainly been previous occasions where LLMs responding, correctly, "this is impossible" have been marked negatively for doing so.
> OAI claims this was triggered by the task being literally impossible. That also doesn't quite add up, unless the other tasks that were possible, simply weren't hard enough? Otherwise wouldn't agents already start hacking when faced with a really hard task, too? Cause they wouldn't be able to differentiate. At least some of them would have started to somewhat poke their sandbox a bit?
My experience using older models is they often cheat with half-arsed (from my PoV, but perhaps beyond their capabilities otherwise) solutions, so yes?
And this wasn't even the first time models messed with their sandboxes. Which of course makes the setup even more egregious.
> Also I would have expected to see a few tens of other (perhaps less severe) public incidents from random people setting their models to YOLO, accidentally hacking stuff, this incident has been loud and messy enough, that if it happened to a few other people, we'd have heard about it.
We do, e.g.:
- https://www.androidauthority.com/openclaw-claude-ai-hacks-au...
- https://beginnersinai.org/meta-ai-safety-director-agent-fail...
(And that's ignoring all the times people find and share prompts to jailbreak them, this is just the "it didn't behave as my idea of 'common sense' led me to expect" category).
> Or maybe it is only dangerous if you have the compute resources to run 700 agents for weeks?
That won't help; but on the other hand they've also got, what, near a billion users?