Comment by danielmarkbruce

2 days ago

You are conflating "half built" with "a piece of a system".

The model weights change as the model goes through the training process. They aren't stored after pre-training is done and other weights are put somewhere else. It's more like pottery - the thing changes. It's not correct to say something is soft and malleable because it once was.

> The model weights change as the model goes through the training process.

Yes. They do. You are absolutely right about that.

But the model architecture doesn't change as a result of the training process. A piston doesn't suddenly turn into a digital watch as a result of tuning an engine. Similarly, the transformer part of a GPT model doesn't suddenly turn into something else as a result of optimizing a loss function.

---

i've got other stuff to do, so i'm stopping here.

  • No one is arguing about the architecture of the model. It's the objective function and optimizer.

    • Just skimming through here but I think you have the wrong ideas with llms, I’d recommend Andrew Ngs course (correct me if you’ve already seen it or something similar).

      8 replies →

You're using the fact the both parts of training affect the same weights to support your argument that they're making the system do something fundamentally different after RL?

  • Assuming you are saying that RL is changing the model from doing one thing to another, yes. RL is changing the nature of the model.