Comment by stusmall

5 days ago

I can't wrap my head around the idea that distillation is IP theft but mass training on books, music and art without consent is fair use. The two stances are incompatible. If it is transformative use of a book, it is transformative use of AI output.

Yeah people in the US who want protectionism for US models on grounds of IP rights are appalling hypocrites. What’s good for the goose is good for the gander!

I really hope that when a sane administration returns to power in the US we'll actually get a reckoning over how irrational it was to not restrict training data.

  • It isn't irrational.

    The idea is that selling, or giving away art grants the consumer to sell art of his own, not a clear copy of.

    So the court rationally found Anthropic guilty for acquiring copies without paying for them. But didn't charge the training, deemed fair use.

    If it isn't fair use, then we may need to sue all teachers for spitting out knowledge they ultimately acquired from someone's work.

    Copyright, some would say is irrational. Humans learn, that is copying, distilling in fact.

    Copyright though is pragmatic: it draws a line, art will be reproduced, let's just enforce that they can't be shameless (near) identical copies. Mix it up, derive the original enough so that you aren't competing with the original author.

    The only argument to ban training on copyright data is that it unfairly compete with original authors. Which stance do you take? Neither would be irrational.

    • I agree that the fair use argument on training data is a genie that’s not getting put back in a bottle. But:

      > If it isn't fair use, then we may need to sue all teachers for spitting out knowledge they ultimately acquired from someone's work.

      Is such a lazy argument and always was. A human doing something is necessarily different than a machine doing it. We can be ok with a human doing a thing and simultaneously not ok with a machine doing it.

      1 reply →

No, it is stealing the R&D of another company.

  • Ok, but only if you also conceed that all of these AI companies were stealing to train their models in the first place.

    Either everyone was stealing all along, and they should all be sued out of existence, or nobody is stealing.

    Personally, I think its all fair use and support all of the training, including distilations.

    • Distillation is forbidden by the provider of the model.

      I suppose you think that model providers are not allowed to impose restrictions on the use of their model? Think carefully before answering.

      3 replies →

Is it? Turning a book into an AI model is lot more transformative than turning an AI Model into another AI Model.