← Back to context

Comment by gcr

12 hours ago

This is oversimplified. The proposed question is whether a model with fewer parameters could achieve performance on one language similar to that of a larger model that’s been trained more broadly, which isn’t straightforward to do.

Oh, I didn’t interpret the above question as asking in that direction; but yeah, that’s of course something I didn’t attempt to answer with my comment.

Although I’d be intrigued in the answer to that small-narrow vs. large-broad model question, too!