← Back to context

Comment by davidos

2 months ago

one of the authors here, AMA.

What sort of regulations did you run into while gathering training data?

  • The main constraint was licensing rather than regulations like GDPR. To ensure compliance, we restricted ourselves to datasets with permissive licenses.

    The one exception was the Genios dataset (a high-quality German corpus), which we used during pretraining under a separate licensing agreement. Beyond that, the growing availability of high-quality permissive datasets meant we could assemble sufficient training data without compromising on quality.