Comment by Philpax

11 hours ago

The point is not to have byte tokens: it's to have only byte tokens, so that the usual failures of tokenisation (e.g. the number of Rs in strawberry) can be avoided.

Would byte tokens really solve how many r’s? It seems like it would require a level of introspective awareness of the tokens being processed that I don’t think transformers have. But I’m not expert enough to go beyond that intuition, so could easily be wrong.

  • Models generally are capable of counting

    • Sure, but they could count bytes without being byte level tokens.

      Can models introspect about the tokens that were passed to them?