Comment by rectang

2 months ago

License the training corpus and encourage copyright suits against outputs from models trained on unlicensed corpora.

This won't work if the courts decide that training is fair use, which certainly seems the direction they are going.

  • Output is a separate issue from training. Courts will never decide that a identical copy spit out by an LLM is non-infringing simply because it went through an LLM stage. Copyright laundering is wishful thinking by tech folks.