Comment by lucrbvi
1 day ago
There are a lot of open-research on pre-training, post-training and RL data mixtures and sourcing.
I recommend checking papers from Datalogy, Nvidia Nemotron, Ai2 (Ollmo, Tulu, ...) and the recent model from Aleph Alpha if you want to learn more.
All of those are good references. Other folks in the thread are missing distinctions between pre training (~the internet + curated sources) and post training (~instructions and RL)