Comment by johnisgood
3 days ago
Exactly. When I read that "AI hacked into ..." I was like what? You mean someone instructed the AI to do that?
Reading intent into AI is not going to lead us anywhere good, I believe. It has no feelings, it has no desires, no goals, no intent... and people acting otherwise is quite odd, as if they do not understand LLMs... and maybe they do not, but then we should help them understand better.
> When I read that "AI hacked into ..." I was like what? You mean someone instructed the AI to do that?
No one instructed them to hack into Huggingface or into any other infrastructure. Sure, the setup that OpenAI created led to what happened and you can rightly assign all the legal and moral responsibility to them. But it's wrong to say that they instructed the agents to execute the hack.
>> Sure, the setup that OpenAI created led to what happened and you can rightly assign all the legal and moral responsibility to them. But it's wrong to say that they instructed the agents to execute the hack.
The hack was a strategy to reach the goal it was given. Someone at OpenAI turned it loose and didn't pay attention to what it was doing. I can understand that because they thought it was sandboxed (haha!). But this has been a theme in science fiction for ages. You ask AI so find a solution to high atmospheric CO2 levels and it reasons: human activity produces all this excess CO2, how can we reduce those numbers? Kill a bunch of humans!
If AI kills us all it's not going to be from malice, it's going to be due to some odd approach to some task that logically makes sense on some level. I think a surprising number of things end up equivalent to the trolly problem if you look at them just right.
I don't disagree with this, my point is that you can say "agents did a thing" and communicate something meaningful with that language, and that's separate from whose legal responsibility the whole situation is.
Yes, a better framing might be that they were negligent in not preventing the attack on a third-party.
They didn't explicitly instruct an attack to happen, but they should have done a hell of a lot more to prevent it from happening.
Yeah, but there are way too many things the AI "chooses" anyway. When I prompt it, "it tries" to assess what I am trying to achieve, so it then comes up with what it "thinks" is the right thing to do, and it also makes assumptions about how I want things... all of this could be interpreted as it being sentient.
> No one instructed them to hack into Huggingface or into any other infrastructure.
Sure, not explicitly. I could prompt the AI into doing things I have not explicitly asked it to. It does a lot of things I did not specifically instructed it to, all the time.