Comment by killerstorm
25 days ago
Well, I mean you mixed up "fine-tuning" and "reinforcement learning" a bit when describing these options.
Regarding the value of these options, SFT communicates more information to the model being trained, but there's a risk of overfitting. So I'd guess they might use both - do a bit of SFT and then finish with RLAIF.
No comments yet
Contribute on Hacker News ↗