Comment by hirvi74
2 days ago
Sure, I get the gist of the article. I have never liked the reductionist argument that LLMs are nothing more than next-token predictors. By that rational, the human brain is really not that much different. When I am having a conversation with another person, I do not usually have every word I will respond with stored in my limited working memory. My output is often predicted based on the previous word I spoke.
> I do not usually have every word I will respond with stored in my limited working memory. My output is often predicted based on the previous word I spoke.
People don't know exactly the words that they're going to say necessarily, but tend to start with a general concept of what they're trying to communicate and only then try to put together the words (sometimes out of order). LLMs do not begin with any sort of concept they're trying to express. LLMs are simulations that attempt to reproduce what an average person might say while wired up to a huge knowledgebase.
> LLMs do not begin with any sort of concept they're trying to express.
Why do the need to? Considering they are merely tools, I actually appreciate they do not do this. A calculator can compute far better than any human, but I appreciate that calculators are not capable of expressing anything about the computations I request. I want the answer, not a conversation.
> LLMs are simulations that attempt to reproduce what an average person might say while wired up to a huge knowledgebase.
If you will allow me to be simplistic, people -- the soul, the self -- are predominately the aggregated effects of memories and experiences and the ability to retain new memories based on new experiences, no? Consider medical conditions in the dementia family of diseases. As memories fade into the ether, what remains of the self?
Also, people simulate/emulate each other all the time based on what an average, reasonable person might say. People incapable or unwilling to perform such mimicry are often labeled with all kinds of pejorative terms.
> I have never liked the reductionist argument that LLMs are nothing more than next-token predictors.
I have never heard such an argument. Recognition that LLMs are nothing more than next-token predictors does not come from reductionism. It comes from simply knowing how they work e.g. from viewing the inference code.
J.S. Bach said something similar about music and keyboard instruments.
> "There's nothing remarkable about it. All one has to do is hit the right keys at the right time and the instrument plays itself."
My issue is not with fact at face value. My issue is with how the fact is often contextually used in arguments to delegitimize and disparage LLM outputs and LLM users.
Yes, LLMs at a fundamental level are next-token predictors. But in my opinion, LLMs are very useful, imperfect next-token predictors.
There are a lot of wannabe John Henry [1] folks out there. Love LLMs or hate'em, most of those John Henry folks ain't beating these machines on a plethora of tasks.
[1] For those unaware, https://en.wikipedia.org/wiki/John_Henry_(folklore)