← Back to context

Comment by chrisjj

1 day ago

> The question being originally asked is whether "next-token predicton" is the right mental model for an RL-trained model,

Regardless, the statement being challenged here is "still next-token prediction, then".

> and I think the answer is no - not only is it not technically correct

It is correct. RL simply adjusts weights - with no effect beyond an equivalent adjustment to the corpus itself. Hence "next-token predictor" remains accurate.

Yeah - but what is it adjusting weights based on? It's not based on next token ....

And per the focus of this thread, regardless of how accurate it is, why do you find "next token predictor" to be the most useful mental model?

> why do you find "next token predictor" to be the most useful mental model?

To me it is an accurate description of the algorithm. And a sufficient explanation for the behaviour. So I don't need it or anything else as a mental model.

I accept this does not suffice for people who cannot comprehend the huge amount of processing and data the empowers it. Lacking a factual understanding, they reach for any mental model as a kind of superstition.

It is sufficiently advanced technology which to many is indistinguishable from magic. This disguises its limitations and enables its limitless false promotion to the gullible, being the reason it is so dangerous to individuals and society.

  • I agree with this, but I just think that with two types of training LLMs are now a two-trick pony rather than a one-trick one.