Comment by bluegatty
1 month ago
No - distillation is not data inputs.
Raw materials vs. Value add.
They are different things, like ore and metal.
Distillation is a new thing we need to understand, it's probably closer to IP than not.
1 month ago
No - distillation is not data inputs.
Raw materials vs. Value add.
They are different things, like ore and metal.
Distillation is a new thing we need to understand, it's probably closer to IP than not.
The "data inputs" were also, very much, somebody's "value added" IP.
We're talking about things like text people wrote, not some kind of raw data floating out in the ether.
Did I say there was no value add in the inputs?
Ore has value, a different kind of value than the output of the refinery.
distillation has been around for 12 years. it's not new in terms of ML techniques.
https://arxiv.org/abs/1503.02531
although i doubt there has been a legal case over it yet in the context of the legality of stealing shit but IANAL.
Yes, I get that, but it's only now the issues are coming into the commons in a way that industry / society needs the regulatory clarity.
It's completey insane that we still don't know how Open Source would work, that the laws are vague and we're still technically waiting for the courts to decide on cases.
The government should a) legislate and b) create test cases and run them through the courts so that we can have clarity.
> it's only now the issues are coming into the commons in a way that industry / society needs the regulatory clarity.
you mean like the regulatory clarity surrounding stealing shit to make the LLMs in the first place?
> [There is] extensive litigation on the limits of Fair Use to AI development. Currently, we only have 3 first instance decisions out of the 53 cases being tried. It will likely take a decade before we understand how Fair Use applies to any one step in AI training, let alone all.
https://www.britishcopyright.org/wp-content/uploads/BCC-Fair...
> The government should a) legislate and b) create test cases and run them through the courts so that we can have clarity.
if so, it would be nice if they approached the instances of stealing shit chronologically. but that's just my view.
57 replies →
Are you suggesting data input is further from IP than distillation?
That would stun me, but it's a little hard to read.
LLM outputs are not copyrightable or rather the user who generated them owns it.
That entirely settles it and there isn’t much else to say about.
If Anthropic feels that other countries are violating their EULA well they are free to stop doing business with them.
The claim these companies make goes much, much further than that.
They claim LLMs "uncopyright" their inputs. So if I take, say, 50 Mickey Mouse comic books, tell ChatGPT to read them and produce 50 "Buster Beagle" comic books that there is ZERO "copyright contamination" and I own those 50 output comics without Disney having any claims on them whatsoever.
Or if I ask ChatGPT to "make a spreadsheet software like Excel, Sheets, Calc, ..." that, again, there is zero copyright claim possible from these people.
It has not been tested, of course.
Do you think writing books (and Wikipedia articles, and stack overflow articles, and github repos, and, and, and, and ...) is not a value add?? What terrible claim.