← Back to context

Comment by chrisjj

4 days ago

> diminishing these things as 'Next token predictors' seems absurdly reductive.

This shows a deep misunderstanding of the paper's claims, which in no way challenge the established view that these bots are next-token predictors.

Regardless, if all you want is a next-token selector, save your money and roll a die.

> This shows a deep misunderstanding of the paper's claims, which in no way challenge the established view that these bots are next-token predictors.

No, this shows an appreciation of the symbolic richness behind that token 'prediction' which the paper leads on.

> Regardless, if all you want is a next-token selector, save your money and roll a die.

Tell me, where is the emergent symbology guiding that dice?

  • > No, this shows an appreciation of the symbolic richness behind that token 'prediction' which the paper leads on.

    The paper claims no symbolic richness beyond that evident from the undisputed next-token prediction.

    > Tell me, where is the emergent symbology guiding that dice?

    There's none. That's my point.

> which in no way challenge the established view that these bots are next-token predictors.

I mean, of course they are, that's literally what the inference loop does. You can look at the source of your favorite model runner and you'll see exactly that.

What I find misleading about this term is that it focuses attention on the "next token" part and glosses over the "prediction" part as some sort of unspecified "statistical algorithm" - even though this is where most of the work happens and where the interesting questions are.

  • I've seen nothing to suggest it misleads anyone else.

    • There are other next token prediction algorithms such as markov chains or HMMs that also "fit the same interface", but are vastly simpler than what LLMs use.

      I've seen various posters talk about "stochastic parrots" or about how LLMs were just "using very simple word statistics" to get their results, which sounded to me very much as if they thought of LLMs simply as glorified markov chains. I would consider that a big misunderstanding.

      1 reply →