Comment by pfdietz
3 hours ago
> It’s far larger than the resulting model.
Is it? How many different books are we talking about, and how much information is that, after conversion to text and lossless compression? Images, maybe, but text?
3 hours ago
> It’s far larger than the resulting model.
Is it? How many different books are we talking about, and how much information is that, after conversion to text and lossless compression? Images, maybe, but text?
These models are trained on way more than just books. GPT-3 was trained on about half a terabyte of filtered plaintext and the training corpuses have grown significantly by then by all accounts.
I imagine that compresses by ~90%, and current top commercial models have a couple of trillion parameters, don't they?
They aren’t trained on compressed plaintext so I’m not sure of the relevance there. But regardless it’s my understanding that’s modern models are trained with orders of magnitude more storage than their parameters require. But it’s possible I’m incorrect. This is getting to the fringe of my knowledge of concrete LLM details.
1 reply →