Comment by gruez
16 days ago
> There’s a mile of difference between citing a work and taking it, rewording it, and not crediting or compensating the original author.
So what's wikipedia doing vs what LLMs do? So far as I can tell the only difference is in citations, but:
1. LLMs can be made to cite, eg. if you use google search's AI mode it'll happily provide citations. I doubt that would placate the AI haters though.
2. Outside of academia no one really cares about citations. There's no legal requirement to cite, nor do I think all the people complaining about AI "stealing" other peoples' work are going to be magically placated by the addition of a few citations. Moreover it's unclear whether the concept of citations makes sense in many contexts. If you ask a human programmer how to write fizzbuzz, they'll likely blurt out a solution without providing citations, much like an AI would. Same for most questions people are asking AI about, eg. "gimme a cake recipe", do you really need a citation back to some 18th century cook book?
> So what's wikipedia doing vs what LLMs do?
For Wikipedia, people look up sources to create some original work. They do it while citing the original work but that's not the major point. That's clearly all lawful, you are allowed to do that.
For LLMs, people (or let a software do it, doesn't really matter) copy original work and use it for commercial purpose. If that's allowed fully depends on the license of the original work, but I didn't think anybody really beliefs that OpenAI and others check the license for every single original work they ever used. So it's basically the biggest copyright infringement ever.
Everything else is just framing coming from big companies.
So, the problem already starts while training the model.
Regarding its output, if it happens to output work that falls under copyright, the LLM company must make sure that it obeys the license connected to it (i.e. citing or not relaying the result to the user). Obviously, nobody does that and it's also not generally possible to do that anyway. So that would be second biggest copyright infringement ever that only works because it's hard to track when such an infringement happens.
So, if asked "could we use your work for our commercial software that might output something that would be still protected by your copyright, but nobody will be able if or when it happens and we won't check and won't tell the users" nobody would have given consent. So they went "duck it, we are talking about billions of dollars and AI is great etcetc., so let's just do it anyway"
I've explained this many times. LLMs don't "copy original work." ChatGPT doesn't contain copies of books inside it, any more than your brain is a copy of the various books you have read. The LLM model isn't physically big enough for that, it's like saying you fit 1000 gallons of water in a 1 gallon milk bottle. Can't be done. The model may be able to quote small snippets of works, just like you might remember various quotes from Shakespeare or someone. (That's not infringing either, by the way)
Wikipedia, and LLMs, can refuse to cite sources and still not infringe copyright. Citation simply isn't relevant to copyright. Not citing a work that you read previously is not a copyright infringement. Wikipedia or OpenAI being non-profit, or for-profit businesses, has nothing to do with copyright. Consent has nothing to do with copyright. Copyrighting something doesn't mean that you can require everyone who reads it to get your consent. You can require everyone who distributes it to get your consent, but once it's been distributed to someone, they can read it freely.
Hope this helps!
>For Wikipedia, people look up sources to create some original work. They do it while citing the original work but that's not the major point. That's clearly all lawful, you are allowed to do that.
>For LLMs, people (or let a software do it, doesn't really matter) copy original work and use it for commercial purpose. If that's allowed fully depends on the license of the original work, but I didn't think anybody really beliefs that OpenAI and others check the license for every single original work they ever used. So it's basically the biggest copyright infringement ever.
So what makes wikipedia (and other encyclopedias) legal but chatgpt not legal? By your own admission citation isn't "the major point". Wikipedia might get a pass because it's a non-profit, but every other encyclopedias operate on the same model.