Comment by angoragoats
2 days ago
> It doesn't check a "prediction" against what was actually "next" in data
Literally no one here is claiming that it does. This is one of the many flaws in the article.
2 days ago
> It doesn't check a "prediction" against what was actually "next" in data
Literally no one here is claiming that it does. This is one of the many flaws in the article.
It does in pre training, but not in RL post training. And not at inference time. Reading over all these comments I get the feeling my mistake was not clearly delineating inference time and train time.
Your mistake was assuming people would be bothered to understand the details of how things work. Most people are lazy and don't know the details of how anything works.
My point is that “it doesn’t check the accuracy of the prediction against the data” is a non-response, because no one calling it a “next-token predictor” is making the claim that it does do that or that they’re calling it a next-token predictor because it does that.
Many people are in fact claiming the thing you are saying they are not - even if you are not. The reason they are claiming it is that it was true at one point, and most intro courses/blog posts/videos still describe them that way and then hand wave some "other stuff at the end". You can even see a comment here that refers to the gpt-2 paper. LLMs were trained to predict the next token, produced a distribution to do so, were scored against their prediction v the truth, and the weights updated so that the probability distribution made it more likely to predict the truth from that sample next time. They were, in every sense of the word, a next token predictor.
They are no longer that thing due to post training. They simply aren't making a prediction, and they aren't even optimized for the next token. If I give a distribution of the heights of the population, I'm not giving a prediction either. Distributions don't imply predictions.
Why the desperation to hang onto the word "prediction"?
1 reply →