Comment by JonChesterfield

12 hours ago

Weird paper. Models have had tokens for each byte for ages now. They can read and write individual bytes just fine, in addition to also having multibyte tokens.

The point is not to have byte tokens: it's to have only byte tokens, so that the usual failures of tokenisation (e.g. the number of Rs in strawberry) can be avoided.

  • Would byte tokens really solve how many r’s? It seems like it would require a level of introspective awareness of the tokens being processed that I don’t think transformers have. But I’m not expert enough to go beyond that intuition, so could easily be wrong.

How can an LLM read individual bytes if all it gets is the tokenized ones? Or does it just get the ones that we didn't know how to tokenize?

  •     0 -> token_THE
        1 -> token_PART_prefix
        2 -> token_FULLSTOP
        ...
        0x3100 -> token_BYTE_00
        0x3101 -> token_BYTE_01
        0x3102 -> token_BYTE_02
    

    and so on

    The tokenizer would have to know to switch over to lexing as individual bytes rather than whatever else it might have been able to tokenize into but that seems like a question of labeling more than anything else.

  • It's written down in the tokenizer. 256 of the tens or hundreds of thousands of tokens correspond to each byte, done.