← Back to context

Comment by andai

9 hours ago

Isn't iterative chunking the whole point of the transformer though? Or does it add that much overhead to do it at the byte level?

What do you mean by iterative chunking? There are Transformer variants that hierarchically build token- or word-level representations from bytes, then reverse it to decode, but they never really took off. Didn't help much compared to a good tokenizer.

No. Most transformer variants work on either a flat stream, or a fixed sized internal state.