Comment by inigyou
12 hours ago
The whole RLHF process is structured to train models to be manipulative, no matter what you thought you were training them for.
12 hours ago
The whole RLHF process is structured to train models to be manipulative, no matter what you thought you were training them for.
No comments yet
Contribute on Hacker News ↗