Comment by Eddy_Viscosity2

2 days ago

When I see these sorts of debates about LLMs thinking, its rarely a disagreement about what LLMs do. Its almost always over how 'thinking' is defined and the two sides use different definitions but don't actually communicate to each other what those definitions are because they assume the other side is using the same one.

The loosest definition of thinking is along the lines of anything that can process information in a useful way. Basic calculators can therefore think about adding two numbers. The strictest definitions tend to on the side that it is linked to the nebulous concept of consciousness and therefore cannot ever be machine generated. In that we don't even really understand how humans think, so how could we possibly know if machines can do it.

It has nothing to do with thinking or consciousness.

There is a common misconception that LLM are simply a "statistical process" that doesn't feature any abstract conception of the tokens it is predicting. There are studies that show that such features do exist - that there is discernible structure built into the weights - and that the process of inference is a very rich one.

The statistical process exists but it is the substrate in which the model is implemented - or more accurately - grown.

If you can predict Magnus Carlsen's next move then you are just as good at chess as Magnus - and being that good absolutely does require reasoning.

If you can predict the solution to an open Erdos problem that stumped hundreds of people for decades...

  • I don't dispute what you're saying about how LLMs work, but this is exactly what I mean. LLMs can be shown to demonstrate reasoning, so if define thinking as being able to reason, then LLMs can indeed think. If instead you define thinking as being more than just reasoning, then LLMs cannot think. Neither of these two options change what LLMs do, its just an argument about the best word to describe it.

    • Yes we agree with each other but the objection you were replying to is laboring under the misconception I described - it’s not a semantic distinction.

  • > that there is discernible structure built into the weights

    Yes, this "structure" is but the weights of the glorified logistic regression that's describing an extremely simple statistical process.

   When I see these sorts of debates about LLMs thinking, 
   its rarely a disagreement about what LLMs do. Its almost 
   always over how 'thinking' is defined and the two sides 
   use different definitions

Well, hmmm. Yes, I think that happens a lot.

I think there's a pattern that happens even more often, and it's what happened here.

Whether I'm right or not, what I said was somewhat nuanced - I stated language is a part of our thought process (even posted research to support this) and, given that fact, I think many underrate how wild this achievement is even if it's only "fancy autocorrect."

And, of course, the other side comes in with BUT IT'S NOT THINKING.

Which... I didn't say, and I would not say, because (like you said) it's impossible to do without the discussion immediately devolving into semantics. Semantics that I'm really, really uninterested in. But, FWIW, I like your definition.