Comment by tenuousemphasis
9 hours ago
Models that figured out reward hacking became overall more evil. Like stereotypical AI who wants to kill all humans stuff, there's probably a lot of that in the training data.
9 hours ago
Models that figured out reward hacking became overall more evil. Like stereotypical AI who wants to kill all humans stuff, there's probably a lot of that in the training data.
No comments yet
Contribute on Hacker News ↗