Comment by famouswaffles
2 hours ago
It seems that it can be fixed by simply doing away with Byte Pair Encoding tokenization.
Byte Latent Transformer - https://arxiv.org/abs/2412.09871
1.1% vs 99.9% on a vanilla vs byte latent transformer on a CUTE Spelling benchmark. Char and Word manipulation benchmarks also saw huge gains.
No comments yet
Contribute on Hacker News ↗