Comment by ivo-42
6 hours ago
I worked on Kolibri, in particular pre-training data and mid-training. We strive to be as open as possible. Glad you like it.
6 hours ago
I worked on Kolibri, in particular pre-training data and mid-training. We strive to be as open as possible. Glad you like it.
How do you cleanse the data at this scale?
By various forms of deduplication (exact, fuzzy, substring), heuristic filters and distilling quality classifiers that annotate our data. Synthetic rephrases can also be considered a form of cleaning/getting more out of existing noisy data.
We have a lot of details in the tech report if you want to go deeper.
Hey! I’m curious if you tried comparing luxical to model2vec classifiers for the pretraining.
I’m one of the authors of model2vec, and working on training classifiers for this. I think model2vec could be better, but I haven’t had the opportunity to try this at scale. So if you did, knowing about it would be helpful!
Are there plans to make much larger versions of this model? With 500B-1T params for general purpose knowledge tasks, similar to the current leading proprietary models?
2 replies →
How much should German authors be paid, $3,000 per book, like Anthropic paid?
Why expose yourself to this liability?