← Back to context

Comment by andai

9 hours ago

How can an LLM read individual bytes if all it gets is the tokenized ones? Or does it just get the ones that we didn't know how to tokenize?

    0 -> token_THE
    1 -> token_PART_prefix
    2 -> token_FULLSTOP
    ...
    0x3100 -> token_BYTE_00
    0x3101 -> token_BYTE_01
    0x3102 -> token_BYTE_02

and so on

The tokenizer would have to know to switch over to lexing as individual bytes rather than whatever else it might have been able to tokenize into but that seems like a question of labeling more than anything else.

It's written down in the tokenizer. 256 of the tens or hundreds of thousands of tokens correspond to each byte, done.