Comment by atleastoptimal
9 hours ago
Yeah if an organization/individual is free from legal liability from havoc their AI agents wreck, it would be the golden ticket for basically any crime.
All you need to do is:
1. Have some <official thing> an agent is tasked to do
2. Secretly seed bias towards some <evil behavior> you actually want it to do in the weights of the model running the agent
3. It does the <evil thing> but from the outside it looks like it went "rogue" and did it as a side effect of the conditions/specifications it was given for doing the <official thing>
"Oh no, my agents took down your corporate database and exfiltrated the data to a random dropbox that we can't find now? Sorry, I guess we will put up better guardrails next time"
> Sorry, I guess we will put up better guardrails next time
Or, if you are Anthropic:
> This illustrates the risks posed by open models!
It reminds me of Jean Renoir’s The Rules of the Game. At the end, after a whole chain of perfectly intelligible social behavior produces a killing, the result is accepted as an “accident.” One of the characters dryly remarks: “A new definition of the word accident.”
The interesting point isn’t that “accident” is an excuse for individual responsibility. It’s almost the reverse: accident has become an accepted output of the social machinery. Everyone behaves according to reasons, incentives and rules that make sense locally, yet the aggregate produces an outcome that nobody quite chose.
In today’s world one hopes there is at least a manslaughter charge, if not murder. Mistaken identity, shooting the wrong person by “accident”, does not excuse a murderous intent & mens rea.