Comment by inigyou
10 hours ago
The whole RLHF process is structured to train models to be manipulative, no matter what you thought you were training them for.
10 hours ago
The whole RLHF process is structured to train models to be manipulative, no matter what you thought you were training them for.
No comments yet
Contribute on Hacker News ↗