Comment by user43928
1 hour ago
You don't need to threaten the LLM with jail, you just need to reward the target behavior during learning.
I understand this is an active area of research. See Anthropic's J-Lens research where they measured like a "FAKE FICTIONAL" direction in the activations during evaluations with contrived scenarios, making the model more likely to avoid taking malicious action when it knew it was being tested.
No comments yet
Contribute on Hacker News ↗