Comment by Philpax
11 hours ago
The point is not to have byte tokens: it's to have only byte tokens, so that the usual failures of tokenisation (e.g. the number of Rs in strawberry) can be avoided.
11 hours ago
The point is not to have byte tokens: it's to have only byte tokens, so that the usual failures of tokenisation (e.g. the number of Rs in strawberry) can be avoided.
Would byte tokens really solve how many r’s? It seems like it would require a level of introspective awareness of the tokens being processed that I don’t think transformers have. But I’m not expert enough to go beyond that intuition, so could easily be wrong.
Models generally are capable of counting
Sure, but they could count bytes without being byte level tokens.
Can models introspect about the tokens that were passed to them?