← Back to context

Comment by jujube3

16 days ago

Yes, Wikipedia gets its content largely by hovering up the web, without the consent of the authors. For example you can cite a New York Times article in Wikipedia, without getting the consent of the NYT author. Wikipedia also "hoovers up" (as you put it) offline sources like books. Again without consent!

Indeed, the need to get consent from the author before reading a published work Isn't A Thing in general, outside some very specific contractural scenarios.

Why are you claiming that citations are the same thing as theft without citation? There’s a mile of difference between citing a work and taking it, rewording it, and not crediting or compensating the original author.

  • The whole job of a Wikipedia author is to take source material and reword and summarize it into an article. They're essentially acting as human LLMs. Wikipedia even has very specific rules against "original research." You are supposed to be acting as a summarizer, not a researcher or author. Wikipedia also does not compensate the original authors of the source material.

    Really the only difference between the Wikipedia author and the LLM is that the Wikipedia author will more frequently be asked to provide citations. But the LLM can also provide citations if asked. In neither case are the authors of what is being summarized compensated or asked for permission. In neither case is it theft.

  • > There’s a mile of difference between citing a work and taking it, rewording it, and not crediting or compensating the original author.

    So what's wikipedia doing vs what LLMs do? So far as I can tell the only difference is in citations, but:

    1. LLMs can be made to cite, eg. if you use google search's AI mode it'll happily provide citations. I doubt that would placate the AI haters though.

    2. Outside of academia no one really cares about citations. There's no legal requirement to cite, nor do I think all the people complaining about AI "stealing" other peoples' work are going to be magically placated by the addition of a few citations. Moreover it's unclear whether the concept of citations makes sense in many contexts. If you ask a human programmer how to write fizzbuzz, they'll likely blurt out a solution without providing citations, much like an AI would. Same for most questions people are asking AI about, eg. "gimme a cake recipe", do you really need a citation back to some 18th century cook book?

    • > So what's wikipedia doing vs what LLMs do?

      For Wikipedia, people look up sources to create some original work. They do it while citing the original work but that's not the major point. That's clearly all lawful, you are allowed to do that.

      For LLMs, people (or let a software do it, doesn't really matter) copy original work and use it for commercial purpose. If that's allowed fully depends on the license of the original work, but I didn't think anybody really beliefs that OpenAI and others check the license for every single original work they ever used. So it's basically the biggest copyright infringement ever.

      Everything else is just framing coming from big companies.

      So, the problem already starts while training the model.

      Regarding its output, if it happens to output work that falls under copyright, the LLM company must make sure that it obeys the license connected to it (i.e. citing or not relaying the result to the user). Obviously, nobody does that and it's also not generally possible to do that anyway. So that would be second biggest copyright infringement ever that only works because it's hard to track when such an infringement happens.

      So, if asked "could we use your work for our commercial software that might output something that would be still protected by your copyright, but nobody will be able if or when it happens and we won't check and won't tell the users" nobody would have given consent. So they went "duck it, we are talking about billions of dollars and AI is great etcetc., so let's just do it anyway"

      2 replies →