← Back to context

Comment by ivo-42

6 hours ago

I worked on Kolibri, in particular pre-training data and mid-training. We strive to be as open as possible. Glad you like it.

How do you cleanse the data at this scale?

  • By various forms of deduplication (exact, fuzzy, substring), heuristic filters and distilling quality classifiers that annotate our data. Synthetic rephrases can also be considered a form of cleaning/getting more out of existing noisy data.

    We have a lot of details in the tech report if you want to go deeper.

    • Hey! I’m curious if you tried comparing luxical to model2vec classifiers for the pretraining.

      I’m one of the authors of model2vec, and working on training classifiers for this. I think model2vec could be better, but I haven’t had the opportunity to try this at scale. So if you did, knowing about it would be helpful!

    • Are there plans to make much larger versions of this model? With 500B-1T params for general purpose knowledge tasks, similar to the current leading proprietary models?

      2 replies →

How much should German authors be paid, $3,000 per book, like Anthropic paid?

Why expose yourself to this liability?