Comment by ericd
1 hour ago
This seems to imply that training will ever be done? But yeah, I think the idea is that the appetite for thinking-on-tap will be enormous.
Even for smaller models, I think they’ve found that training an enormous, inefficient model and then distilling it internally to something much more efficient to serve is the way to go.
No comments yet
Contribute on Hacker News ↗