Comment by estearum
10 hours ago
It's not even fair to call "run code" to be indirect compared to what a human would do. The word raspberry has no Rs in it in human language either. We have a written representation of it, which we can then write down either in our head or on paper, and then we can "run the algorithm" of counting each of the letters.
Nothing intrinsically more or less direct about the LLM's method than ours.
I could argue LLM only have "token" as their perceivable dimension, compare to human multiple senses as the physic perceivable dimension and a brain with many other dimension of "learning" and "thinking". In spoken language, we may not have 'r' but in written we have, both spoken language and written language are learned skills.
You could argue in return that humans only have electro-chemistry as our one perceivable dimension. We only indirectly perceive light through the signals our eyes send to our brains.
In my mind general intelligence is pretty much by definition a virtual machine, so the mechanisms behind thought are only relevant for the sake of efficiency (ie you can argue that LLMs make a poor basis for intelligence because tokens and natural language are a poor way to encode the world, but if you can run it on a big enough computer to counteract the inherent wasteful virtualisation then who really cares how it works under the hood?)
So LLM and human all have 1 dimenion perceivable signal, just LLM is 240p, and human is 8K in resolution, that's why we have 'r' in our signal, LLM still have 'r' in their signal, just because of the "low resolution", raspberry wasn't encoded with so many 'r' as in human signal.
I will stop here before our analogies go too far.
Is "token" a directly perceivable unit for the LLM? If you ask it "how many tokens are in this sentence?" can it count them (again, not guessing or making a tool call)?
I've never tried it and it might take some thought and effort to conduct an experiment to find out properly, but I would be interested in the answer.
I dont think so. This is akin to asking a person, what is the frequency of the light hitting your eye when watching a leaf for example. You either know the (approximate) answer by knowing the frequency of green, or use a tool to measure it. If the LLM gives the correct answer it is either.guessing based on intution(and this intuition is based on pairs of word to tokenization length in text form in training data), writing code(or executing a tokenizer) or running a tokenizer mentally (reasoning via CoT).
Not the point: the simulated intelligence in this context needs to create proper representation. It is not a matter of what it sees but of what it can see.