Comment by bulbar
16 days ago
> So what's wikipedia doing vs what LLMs do?
For Wikipedia, people look up sources to create some original work. They do it while citing the original work but that's not the major point. That's clearly all lawful, you are allowed to do that.
For LLMs, people (or let a software do it, doesn't really matter) copy original work and use it for commercial purpose. If that's allowed fully depends on the license of the original work, but I didn't think anybody really beliefs that OpenAI and others check the license for every single original work they ever used. So it's basically the biggest copyright infringement ever.
Everything else is just framing coming from big companies.
So, the problem already starts while training the model.
Regarding its output, if it happens to output work that falls under copyright, the LLM company must make sure that it obeys the license connected to it (i.e. citing or not relaying the result to the user). Obviously, nobody does that and it's also not generally possible to do that anyway. So that would be second biggest copyright infringement ever that only works because it's hard to track when such an infringement happens.
So, if asked "could we use your work for our commercial software that might output something that would be still protected by your copyright, but nobody will be able if or when it happens and we won't check and won't tell the users" nobody would have given consent. So they went "duck it, we are talking about billions of dollars and AI is great etcetc., so let's just do it anyway"
I've explained this many times. LLMs don't "copy original work." ChatGPT doesn't contain copies of books inside it, any more than your brain is a copy of the various books you have read. The LLM model isn't physically big enough for that, it's like saying you fit 1000 gallons of water in a 1 gallon milk bottle. Can't be done. The model may be able to quote small snippets of works, just like you might remember various quotes from Shakespeare or someone. (That's not infringing either, by the way)
Wikipedia, and LLMs, can refuse to cite sources and still not infringe copyright. Citation simply isn't relevant to copyright. Not citing a work that you read previously is not a copyright infringement. Wikipedia or OpenAI being non-profit, or for-profit businesses, has nothing to do with copyright. Consent has nothing to do with copyright. Copyrighting something doesn't mean that you can require everyone who reads it to get your consent. You can require everyone who distributes it to get your consent, but once it's been distributed to someone, they can read it freely.
Hope this helps!
>For Wikipedia, people look up sources to create some original work. They do it while citing the original work but that's not the major point. That's clearly all lawful, you are allowed to do that.
>For LLMs, people (or let a software do it, doesn't really matter) copy original work and use it for commercial purpose. If that's allowed fully depends on the license of the original work, but I didn't think anybody really beliefs that OpenAI and others check the license for every single original work they ever used. So it's basically the biggest copyright infringement ever.
So what makes wikipedia (and other encyclopedias) legal but chatgpt not legal? By your own admission citation isn't "the major point". Wikipedia might get a pass because it's a non-profit, but every other encyclopedias operate on the same model.