Comment by reissbaker
20 hours ago
Instead of yet another mediocre but fully-made-in-the-West open model (alongside Mistral, Trinity, Poolside, Inkling, etc etc) I'd really love for a Western neloab start the same way Qwen did: by focusing on post-training. Qwen's first release was a Llama 1 finetune [1]! Once they made it useful, they started working their way back in the stack to also do their own pretraining, etc. Starting with pretraining feels like such a waste: there's millions of dollars of crystallized compute and data sitting around in the Chinese model weights. Why not start with one of those, and only work your way back to pretraining once you've released something you can prove is useful?
I think there aren't any recent base models to post train on. Labs don't release them any more.
At least for RL, you don't need a base model — the rollouts are run in an inference engine with an instruction-tuned model using a chat template. You can start with just that!
Yeah but they are already RL'd too. No doubt you can further improve an RL'd model by doing your own better RL on top of that. But it's not going to be the same as starting with a base model or instruction tuned model.
https://huggingface.co/IFM
This is point of Nemotron series