Comment by armcat
14 hours ago
There has been extensive research into token-free LLMs, but for some reason we are still operating in a token domain, so there is something to that.
Byte Latent Transformers (BLT): https://arxiv.org/abs/2412.09871
Charformer: https://arxiv.org/abs/2106.12672
It's because of "chunking" which human minds do as well. If you increase granularity, you increase permutations and required compute to match the same performance of existing systems.
GPT4's tokeniser for example made use of extended tokens for coding structures and with enough training on code was quite ahead of the rest for a while.
The tokeniser and its vocabulary makes a big difference.
Isn't iterative chunking the whole point of the transformer though? Or does it add that much overhead to do it at the byte level?
What do you mean by iterative chunking? There are Transformer variants that hierarchically build token- or word-level representations from bytes, then reverse it to decode, but they never really took off. Didn't help much compared to a good tokenizer.
No. Most transformer variants work on either a flat stream, or a fixed sized internal state.
Aren't diffusion models token-free?
No, they still operate on tokens, it's just the difference in how text is sampled. It's done through iterative denoising and in parallel, across multiple token positions.
H-Net: https://proceedings.iclr.cc/paper_files/paper/2026/hash/f1aa...