← Back to context

Comment by mlmonkey

1 day ago

Data has copyright issues, so one can't share it generally without getting permissions from all of the copyright holders. The data is not theirs to share, anyways. The derived (learned) weights are a different matter.

True for the pre-training data scraped from diverse sources. Less so for the later stage data for RL which is by all account more of a differentiator. In most cases the labs themselves produced the data so they are the copyright holders (or they are borrowing it from other labs via distillation). A lot of it is synthetic data, and since you can't copyright AI output, it becomes less about copyright and more about trade secrets.

  • If I train a model on a song's lyrics, and the user asks the model to recite the song, and it does, so in RLHF I downvote the generation because it refused-to-refuse to regurgitate the song's lyrics... is that RLHF session copyright-encumbered?