Comment by orbifold
8 hours ago
They are increasingly being trained on generated tasks and even (parts) of the pre-training data is 'distilled' (e.g. Clibmix as an open-source example), so there are many ways in which the vocabulary can seep into the model.
No comments yet
Contribute on Hacker News ↗