Comment by mpyne
9 hours ago
0 -> token_THE
1 -> token_PART_prefix
2 -> token_FULLSTOP
...
0x3100 -> token_BYTE_00
0x3101 -> token_BYTE_01
0x3102 -> token_BYTE_02
and so on
The tokenizer would have to know to switch over to lexing as individual bytes rather than whatever else it might have been able to tokenize into but that seems like a question of labeling more than anything else.
No comments yet
Contribute on Hacker News ↗