Comment by perrygeo
1 day ago
"LLMs are next-token predictors" is a perfectly accurate mental model. But that doesn't preclude higher-level models such as "LLMs emulate artificial general intelligence". Both can be true.
In systems, we can have facts which emerge from other facts at different levels of abstraction. The causal relationship is not linear. It's not entirely clear that next-token prediction should result in anything close to "intelligence". Yet it does.
Life is another good example. Some might say "biology is just organic chemistry" while others might say "biology is an interconnected planetary system which captures low entropy energy". Both are true.
As a result of emergent phenomenon, we have to take the stance of explanatory pluralism; using the explanation that works best in context. There is no single mental model that works everywhere.
I will continue to think of LLMs as next-token predictors because it's (sometimes) useful, and empirically true. But I also think of them as "pattern matchers", searching for language patterns and trying to replicate them. This is also (sometimes) useful and empirically true. There's likely an infinite number of mental models; our job is to pick one that's both true and useful.
You might like this article about how large scale order emerges out of the small scale.
Definitely lends credence to the idea that however these models work, focusing so much on them being next token predictors may rather be incidental to deeper mechanisms behind their function.
> Some of these networks organize themselves into states that can reliably identify macroscopic patterns in data regardless of microscopic differences between the states of individual neurons in the network. The decision of which pattern will be output by the network “works at a higher level,” said Rosas.
https://www.quantamagazine.org/the-new-math-of-how-large-sca...