Comment by ordersofmag
1 day ago
True for the pre-training data scraped from diverse sources. Less so for the later stage data for RL which is by all account more of a differentiator. In most cases the labs themselves produced the data so they are the copyright holders (or they are borrowing it from other labs via distillation). A lot of it is synthetic data, and since you can't copyright AI output, it becomes less about copyright and more about trade secrets.
If I train a model on a song's lyrics, and the user asks the model to recite the song, and it does, so in RLHF I downvote the generation because it refused-to-refuse to regurgitate the song's lyrics... is that RLHF session copyright-encumbered?