Comment by immibis
3 days ago
It's not about delivering punishment. It's about suppressing certain responses. If the model is trained seeing that responses using don't contain things that previous messages say will be punished then that is a valid way to deprioritize those responses.
No comments yet
Contribute on Hacker News ↗