It's because of "chunking" which human minds do as well. If you increase granularity, you increase permutations and required compute to match the same performance of existing systems.
GPT4's tokeniser for example made use of extended tokens for coding structures and with enough training on code was quite ahead of the rest for a while.
The tokeniser and its vocabulary makes a big difference.
No, they still operate on tokens, it's just the difference in how text is sampled. It's done through iterative denoising and in parallel, across multiple token positions.
Wow nature got rolled. Should have stayed closer to their expertise. They were already being hustled by a lot of the applied AI work they were accepting.
Obvious work that is behind state of the art. Not field defining. A reasonable paper to publish at NeurIPS or ICML but very middle of the pack.
If your paper is accepted to Nature it should be among the top results in your field for the year. This is just fine.
Edit to clarify: Nature has not been, historically, a venue for pure machine learning papers. It's been a venue for field changing work in the physical sciences. They already have a Machine Intelligence subjournal.
What this paper shows me is how desperate they are for ML papers, and how poorly their staff understand the field.
Weird paper. Models have had tokens for each byte for ages now. They can read and write individual bytes just fine, in addition to also having multibyte tokens.
The point is not to have byte tokens: it's to have only byte tokens, so that the usual failures of tokenisation (e.g. the number of Rs in strawberry) can be avoided.
Would byte tokens really solve how many r’s? It seems like it would require a level of introspective awareness of the tokens being processed that I don’t think transformers have. But I’m not expert enough to go beyond that intuition, so could easily be wrong.
The tokenizer would have to know to switch over to lexing as individual bytes rather than whatever else it might have been able to tokenize into but that seems like a question of labeling more than anything else.
There has been extensive research into token-free LLMs, but for some reason we are still operating in a token domain, so there is something to that.
Byte Latent Transformers (BLT): https://arxiv.org/abs/2412.09871
Charformer: https://arxiv.org/abs/2106.12672
It's because of "chunking" which human minds do as well. If you increase granularity, you increase permutations and required compute to match the same performance of existing systems.
GPT4's tokeniser for example made use of extended tokens for coding structures and with enough training on code was quite ahead of the rest for a while.
The tokeniser and its vocabulary makes a big difference.
Isn't iterative chunking the whole point of the transformer though? Or does it add that much overhead to do it at the byte level?
Aren't diffusion models token-free?
No, they still operate on tokens, it's just the difference in how text is sampled. It's done through iterative denoising and in parallel, across multiple token positions.
H-Net: https://proceedings.iclr.cc/paper_files/paper/2026/hash/f1aa...
I don't understand how this is different from BPE tokenization.
subword tokenizers were never causal in the first place, so BPE was peeking at future bytes all along! TIL
Wow nature got rolled. Should have stayed closer to their expertise. They were already being hustled by a lot of the applied AI work they were accepting.
What’s bad about this paper?
Obvious work that is behind state of the art. Not field defining. A reasonable paper to publish at NeurIPS or ICML but very middle of the pack.
If your paper is accepted to Nature it should be among the top results in your field for the year. This is just fine.
Edit to clarify: Nature has not been, historically, a venue for pure machine learning papers. It's been a venue for field changing work in the physical sciences. They already have a Machine Intelligence subjournal.
What this paper shows me is how desperate they are for ML papers, and how poorly their staff understand the field.
3 replies →
[flagged]
[flagged]
Weird paper. Models have had tokens for each byte for ages now. They can read and write individual bytes just fine, in addition to also having multibyte tokens.
The point is not to have byte tokens: it's to have only byte tokens, so that the usual failures of tokenisation (e.g. the number of Rs in strawberry) can be avoided.
Would byte tokens really solve how many r’s? It seems like it would require a level of introspective awareness of the tokens being processed that I don’t think transformers have. But I’m not expert enough to go beyond that intuition, so could easily be wrong.
How can an LLM read individual bytes if all it gets is the tokenized ones? Or does it just get the ones that we didn't know how to tokenize?
and so on
The tokenizer would have to know to switch over to lexing as individual bytes rather than whatever else it might have been able to tokenize into but that seems like a question of labeling more than anything else.