Comment by amelius

4 days ago

What I want is a model that is trained with data that is openly available, where the data is curated by academia. I don't want corporate crap in my AI (unless it has been filtered properly).

I'm ready to stand corrected, but I'm pretty positive that such a process would require 1) an insane amount of work and 2) wouldn't produce anything close to SOTA results because of the lack of training data.

It is my understanding that - sadly - the insane amount of copyrighted works and corporate crap is a prerequisite for having a corpus that is big enough

  • I think these days even the frontier labs are using large amounts of synthetic data too, which must be worth it even though it seems like an Ouroborous.

There are models that do that, for example the Olmo family of models, although they have Gemma 3 performance levels for that matter.