Comment by peddling-brink
10 hours ago
Humans will and do absolutely do this when there are no consequences.
Humans on a red team, with rules of engagement, that don’t want to go to prison, won’t do this.
We could threaten an LLM with jail, but if it’s sufficiently intelligent, it will realize this is an empty threat. And I’m not sure that building a survival instinct in is going to solve the alignment problem either.
You don't need to threaten the LLM with jail, you just need to reward the target behavior during learning.
I understand this is an active area of research. See Anthropic's J-Lens research where they measured like a "FAKE FICTIONAL" direction in the activations during evaluations with contrived scenarios, making the model more likely to avoid taking malicious action when it knew it was being tested.