Comment by ericd

2 hours ago

This seems to imply that training will ever be done? But yeah, I think the idea is that the appetite for thinking-on-tap will be enormous.

Even for smaller models, I think they’ve found that training an enormous, inefficient model and then distilling it internally to something much more efficient to serve is the way to go.