Comment by otabdeveloper4
2 days ago
> sort of crystallized a bit of the human thought process
a) LLMs don't think. They predict a most probable sequence of language tokens. Huge difference there.
b) Whatever LLMs do doesn't model human behavior whatsoever. LLMs are basically very fancy logistic regressors. I.e., it's a mathematical abstraction first and foremost.
Did you reply to the right post? You quoted me, but you wrote "LLMs don't think" as if it was a rebuttal. It's puzzling, because I didn't say that they think, so it kinda seems like you got confused? Maybe somebody else said that?
I don't really have an opinion on whether or not they "think" because I feel it's impossible to even discuss without getting into a very very uninteresting semantic argument about what "thinking" is.
Are we defining "thinking" as doing it the same way humans do it? Then, of course they're not thinking. It's a statistical model, not axons and neurons, or even a simulation of axons and neurons.
Are we defining "thinking" on a purely functional or behavioral basis, kind of a Turing test approach? Then... well, I think it gets nuanced. For some tasks, within some constraints, they do pass that test. For many others, of course they don't.
Are we defining thinking in more esoteric terms? Something to do with the soul? Maybe the ability to come up with truly novel concepts rather than rehashing and remixing the stuff it was trained on? Do ants think? Do dogs think? Do jellyfish think? Octopi? A newborn baby?
Anyway, it's a deeply uninteresting semantic question.
When I see these sorts of debates about LLMs thinking, its rarely a disagreement about what LLMs do. Its almost always over how 'thinking' is defined and the two sides use different definitions but don't actually communicate to each other what those definitions are because they assume the other side is using the same one.
The loosest definition of thinking is along the lines of anything that can process information in a useful way. Basic calculators can therefore think about adding two numbers. The strictest definitions tend to on the side that it is linked to the nebulous concept of consciousness and therefore cannot ever be machine generated. In that we don't even really understand how humans think, so how could we possibly know if machines can do it.
It has nothing to do with thinking or consciousness.
There is a common misconception that LLM are simply a "statistical process" that doesn't feature any abstract conception of the tokens it is predicting. There are studies that show that such features do exist - that there is discernible structure built into the weights - and that the process of inference is a very rich one.
The statistical process exists but it is the substrate in which the model is implemented - or more accurately - grown.
If you can predict Magnus Carlsen's next move then you are just as good at chess as Magnus - and being that good absolutely does require reasoning.
If you can predict the solution to an open Erdos problem that stumped hundreds of people for decades...
I don't dispute what you're saying about how LLMs work, but this is exactly what I mean. LLMs can be shown to demonstrate reasoning, so if define thinking as being able to reason, then LLMs can indeed think. If instead you define thinking as being more than just reasoning, then LLMs cannot think. Neither of these two options change what LLMs do, its just an argument about the best word to describe it.
1 reply →
> that there is discernible structure built into the weights
Yes, this "structure" is but the weights of the glorified logistic regression that's describing an extremely simple statistical process.
Well, hmmm. Yes, I think that happens a lot.
I think there's a pattern that happens even more often, and it's what happened here.
Whether I'm right or not, what I said was somewhat nuanced - I stated language is a part of our thought process (even posted research to support this) and, given that fact, I think many underrate how wild this achievement is even if it's only "fancy autocorrect."
And, of course, the other side comes in with BUT IT'S NOT THINKING.
Which... I didn't say, and I would not say, because (like you said) it's impossible to do without the discussion immediately devolving into semantics. Semantics that I'm really, really uninterested in. But, FWIW, I like your definition.
The (or a) current neuroscience models of the brain are that its main job is to predict how the body should be responding in the near future. Obviously a lot more complex network nodes than an LLM, but prediction is clearly tied up with thought in some way.
I don’t find LLMs to be very good independent thinkers, but I wouldn’t over sell our own mentation either - it clearly arises from a large number of simpler entities.
The more significant difference is that the LLM is stuck with language which is clearly an emergent and secondary capability of our own thinking. We can formulate words to explain things, but we also can look at two volumes and feel what it means that one is larger than the other. Raise a toddler and you can see the progression from not understanding, repeated experiments, muscle memory and finally to conscious point for reasoning.
Yeah. I don't see them ever hitting the heights of human creativity in terms of coming up with entirely new ideas, schools of thought, etc. That really might be a fundamental limitation of being trained on existing thought. Also, a lot of human experience involves (1) things we don't have words for (2) things we've never put into words.
Yes. And it's part of our thinking. More than a capability . Thought influences speech, but speech also influences thought.
That's why I think it's remarkable that we've managed to (choosing my words very, very carefully here) create a statistical model that does a remarkably decent job at emulating the behavior of a fragment of that process.
It is remarkable. A shame that unregulated business models and giant pools of capital seeking very high returns seem to be distorting (hype cycle, excessive spending that seems untethered to realistic plans, and uncritical roll out prior to compelling proof of utility) the rollout of the remarkable new technology.
I would speculate when we do eventually develop independent synthetic sentient beings, LLM technology will be a part of the package. Perhaps also growing up with a sibling that tries to trick one.
Maybe someone needs to write the singularity novel but with Cain and Abel, not just a unified super intelligence but siblings full of some good will and a good bit of clear seeing and some fun (?) trickery.
The statistical model that underpins any deep learning system is the substrate in which a process is implemented. There is still a process - just because that process is grown - not programmed - doesn't tell us anything about the depth or limits of its capability.
"Fancy logistic regressors" are, in fact, a modelling tool. You can tell LLMs model human behaviour because they're doing things that until a few years ago only humans could do, like cheat in exams.
It's amazing that you can predict a counterexample to an open math problem, all without thinking.
Yet they do.
My favourite quote on this subject... I forgot by who.
"They don't think, they only seem to think. And likewise, they won't replace the majority of human labor, they will only seem to do so."
Recalling "They're Made of Weights"
"If only we had a word for that process that happens before the words come out."
"Artist Formerly Known as Thinking"
Not really, a shitload of "open math problems" are bounded by constraints of simple text processing or heuristic search.
Much of math is just boring routine work.
You should look into emergence. An ant in a colony, an offer in a market, a drop of water in a weather system are all evidence that irreducibly complex things have simple mechanisms at their core.