Comment by cavoirom
10 hours ago
counting 'r' in 'raspberry' to the LLM is similar to 4-dimension space to human. Their world's unit is token, not character, although they could use indirect method such as "run code" to find out. It will stay that way until they change the fundamental of the token that the LLM can perceive characters.
I hope you understand: it is a core point that systems that answer questions must have the ability to internally represent the objects they assess in a way that allows reliability. Whatever the object.
I’m working on this problem using a vocab-free, byte-based approach. It’s definitely solvable.
https://huggingface.co/posts/omarkamali/593639295164067
https://huggingface.co/blog/omarkamali/tokenization
Careful: the problem is very certainly ___not___ counting letters. That is only a telling way to check "is the NN checking or not?". We demand that NNs for consultancy tasks check, strictly.
I used to think byte level tokenization was the answer, but humans also think at a word level and only reevaluate the words at a character level when asked. The solution to better tokenization across languages is likely to be learned tokenization. Here is one attempt I have seen: https://github.com/SamD770/bitter-lesson-tokenization
How many 'r's are there in the next 30 seconds of this [1] song?
[1]: https://youtu.be/l7vRSu_wsNc?si=SndkB6GBaRyhvNNA&t=61
It's not even fair to call "run code" to be indirect compared to what a human would do. The word raspberry has no Rs in it in human language either. We have a written representation of it, which we can then write down either in our head or on paper, and then we can "run the algorithm" of counting each of the letters.
Nothing intrinsically more or less direct about the LLM's method than ours.
I could argue LLM only have "token" as their perceivable dimension, compare to human multiple senses as the physic perceivable dimension and a brain with many other dimension of "learning" and "thinking". In spoken language, we may not have 'r' but in written we have, both spoken language and written language are learned skills.
You could argue in return that humans only have electro-chemistry as our one perceivable dimension. We only indirectly perceive light through the signals our eyes send to our brains.
In my mind general intelligence is pretty much by definition a virtual machine, so the mechanisms behind thought are only relevant for the sake of efficiency (ie you can argue that LLMs make a poor basis for intelligence because tokens and natural language are a poor way to encode the world, but if you can run it on a big enough computer to counteract the inherent wasteful virtualisation then who really cares how it works under the hood?)
1 reply →
Is "token" a directly perceivable unit for the LLM? If you ask it "how many tokens are in this sentence?" can it count them (again, not guessing or making a tool call)?
I've never tried it and it might take some thought and effort to conduct an experiment to find out properly, but I would be interested in the answer.
2 replies →