Comment by danielmarkbruce
1 day ago
You are conflating "half built" with "a piece of a system".
The model weights change as the model goes through the training process. They aren't stored after pre-training is done and other weights are put somewhere else. It's more like pottery - the thing changes. It's not correct to say something is soft and malleable because it once was.
> The model weights change as the model goes through the training process.
Yes. They do. You are absolutely right about that.
But the model architecture doesn't change as a result of the training process. A piston doesn't suddenly turn into a digital watch as a result of tuning an engine. Similarly, the transformer part of a GPT model doesn't suddenly turn into something else as a result of optimizing a loss function.
---
i've got other stuff to do, so i'm stopping here.
No one is arguing about the architecture of the model. It's the objective function and optimizer.
Just skimming through here but I think you have the wrong ideas with llms, I’d recommend Andrew Ngs course (correct me if you’ve already seen it or something similar).
8 replies →
You're using the fact the both parts of training affect the same weights to support your argument that they're making the system do something fundamentally different after RL?
Assuming you are saying that RL is changing the model from doing one thing to another, yes. RL is changing the nature of the model.