Comment by lern_too_spel
6 hours ago
I used to think byte level tokenization was the answer, but humans also think at a word level and only reevaluate the words at a character level when asked. The solution to better tokenization across languages is likely to be learned tokenization. Here is one attempt I have seen: https://github.com/SamD770/bitter-lesson-tokenization
No comments yet
Contribute on Hacker News ↗