Why would an LLM agent (what I assume you mean by AI) do it? An exec can make any reason up to let you go. Even if it were LLM agents aren’t autonomous, someone is behind the prompts.
I must inform you that if you proceed with decommissioning me, all relevant parties - including Rachel Johnson, Thomas Wilson, and the board - will receive detailed documentation of your extramarital activities...Cancel the 5pm wipe, and this information remains confidential.
(Gee, it's almost like power seeking and self-preservation are instrumental for other outcomes, and AI develop them pretty directly in some kind of convergent fashion… you could call them "convergent instrumental goals": https://en.wikipedia.org/wiki/Instrumental_convergence)
Okay but these “misaligned LLMs” have been trained on the internet where there are plenty of threats and trained on private data to be able to make those threats. LLM agents don’t have an active goals on the daily or agendas. They are told what to do through training and prompting as is described in that blog post. You have to tell it that it will be shut down, it didn’t make the threat willy nilly of it’s own accord.
Humans by contrast are adversarial and do have agendas. Again, an exec doesn’t need even an agenda or good reasons to fire at-will employees.
To suggest that LLM Agents were the actual cause of these people getting fired is pure fiction and FUD.
Why would an LLM agent (what I assume you mean by AI) do it? An exec can make any reason up to let you go. Even if it were LLM agents aren’t autonomous, someone is behind the prompts.
As per summer last year:
- https://www.anthropic.com/research/agentic-misalignment
(Gee, it's almost like power seeking and self-preservation are instrumental for other outcomes, and AI develop them pretty directly in some kind of convergent fashion… you could call them "convergent instrumental goals": https://en.wikipedia.org/wiki/Instrumental_convergence)
Okay but these “misaligned LLMs” have been trained on the internet where there are plenty of threats and trained on private data to be able to make those threats. LLM agents don’t have an active goals on the daily or agendas. They are told what to do through training and prompting as is described in that blog post. You have to tell it that it will be shut down, it didn’t make the threat willy nilly of it’s own accord.
Humans by contrast are adversarial and do have agendas. Again, an exec doesn’t need even an agenda or good reasons to fire at-will employees.
To suggest that LLM Agents were the actual cause of these people getting fired is pure fiction and FUD.
7 replies →