Comment by MuffinFlavored
3 years ago
> RLHF
Reinforcement Learning from Human Feedback
Aren't these systems already trained to score good things higher and bad things worse dictated by human feedback?
3 years ago
> RLHF
Reinforcement Learning from Human Feedback
Aren't these systems already trained to score good things higher and bad things worse dictated by human feedback?
personalized RLHF is the keyword