Comment by notpachet
16 hours ago
Related reading:
The Waluigi Effect: After you train an LLM to satisfy a desirable property, then it's easier to elicit the chatbot into satisfying the exact opposite property.
https://www.lesswrong.com/posts/D7PumeYTDPfBTp3i7/the-waluig...
No comments yet
Contribute on Hacker News ↗