Comment by michaelmior

1 year ago

"Reading stuff freely posted on the internet" is also very different from a business having machines consume large volumes of data posted on the Internet for the purpose of generating value for them without compensating the creators. I'm not making a value judgement one way or the other, but "reading stuff freely posted on the Internet" is an oversimplification.

Okay, but "stealing" is also an oversimplification, to the point of absurdity.

It makes no sense to put stuff up on the internet where it can freely be downloaded by anyone at any time, by people who are then free to do whatever they like with it on their own hardware, then complain that people have downloaded that stuff and done what they liked with it on their own hardware.

"Having machines consume large volumes of data posted on the Internet for the purpose of generating value for them without compensating the creators" is equally a description of Google.

  • They are not free to do whatever they like, there are tomes of laws across all countries governing what someone can and cannot do with your intellectual property. Just because we didn't have the foresight to add in a "if by chance in the future someone invents artificial intelligence, that's not fair use" is a shame, but doesn't make what these companies are doing ethical or morale.

    I don't disagree regarding Google, I also think they exploited others IP for their own gain. It was once symbiotic with webmasters, but when that stopped they broke that implied good faith contract. In a sense, their snippets and widgets using others IP and no longer providing traffic to the site was the warning shot for where we are now. We should have been modernising IP laws back then.

    • I did say "free to do whatever they like on their own hardware", because intellectual property laws generally govern the transfer of such property rather than the use.

      After seeing the harm done by the expansion of patent law to cover software algorithms, and the relentless abuse done under the DMCA, I am reflexively skeptical of any effort to expand intellectual property concepts.

      9 replies →

  • What if that data isn’t publicly posted? For example, copilot regurgitating code from private repos, complete with comments.

  • Proper term for it is Computer Assisted Plagiarism, CAP for short. Also, I really hope that Google doesn't claim it created sites it crawl for search their engine.

  • Ok so if I publish under a license saying I don't allow for it to be used for AI do you believe they respect it? What word would you use to describe this violation? Go ahead throw up a robots.txt, throw up a license. You will be able to coax the "fair use" stochastic parrots to render it verbatim.

    Sam Altman and his ilk are exploiting the incredibly slow moving legal system to enrich themselves.

    • > Ok so if I publish under a license saying I don't allow for it to be used for AI do you believe they respect it?

      It's even worse than that, they don't even legally have to respect it if courts find it to be fair use, and so far they have. If it's fair use to train models on it, your license means nothing.

      The only way to "win" is to not publish your code at all, anywhere.

  • It's not about the downloading of the data, it's about its use in training models, which is dubious from a copyright perspective.

  • That is not at all how the internet works. Try to download music from Napster and Lars will sue your ass.

    • No he certainly will not; you will only get sued if you upload Lars' music to share with other people. If you download an illegal copy, the person you downloaded from is the one breaking the law.

      1 reply →

  •   > where it can freely be downloaded by anyone at any time, by people who are then free to do whatever they like with it on their own hardware
    

    I think you have a strong misunderstanding of the law and the general expectation of others.

    I'd like to remind you that a lot of celebrities face legal issues for posting photos of themselves. Here's a recent example with Jennifer Lopez[0]. The reason these types of lawsuits are successful is because it is theft of labor. If you hire a professional photographer to take photos of your wedding then the contract is that the photographer is handing over ownership of the photos in exchange of payment. The only difference here is that the photo was taken before a contract was made. The celebrity owns the right to their body and image, but not to the photograph.

    Or think about Open Source Software. Just because it is posted on GitHub does not mean you are legally allowed to use it indiscriminately. GitHub has licenses and not all of them are unrestricted. In fact, a repo without a license does not mean unfettered usage. The default is that the repo owner has the copyright[1].

      > You're under no obligation to choose a license. However, without a license, the default copyright laws apply, meaning that you retain all rights to your source code and no one may reproduce, distribute, or create derivative works from your work.
    

    A big part of what will make a lawsuit successful or not is if the owner has been deprived of compensation. As in, if you make money off of someone else's work. That's why this has been the key issue in all these AI lawsuits. Where the question is about if the work is transformative or not. All of this is in new legal territory because the laws were not written with this usage in mind. The transformative stuff is because you need to allow for parody or referencing. You don't want a situation where, say... someone including a video of what the president has said to discuss what was said[2]. But this situation is much closer to "Joe stole a book, learned from that book, and made a lot of money through the knowledge that they obtained from this book AND would not have been able to do without the book's help." Just, it's usually easier to go after the theft part of that situation. It's definitely a messy space.

    But basically, just because a piece of art exists on public property does not mean you have the right to do whatever you want with it.

      >  is equally a description of Google.
    

    Yes and no. The AI summaries? Yeah. The search engine and linking? No. The latter is a mutually beneficial service. It's one thing to own a taxi service and it is another to offer a taxi service that will walk into a starbucks take a random drink off the counter and deliver it to you. I'm not sure why this is difficult to understand.

    [0] https://www.bbc.com/news/articles/cx2qqew643go

    [1] https://docs.github.com/en/repositories/managing-your-reposi...

    [2] https://www.youtube.com/watch?v=tUnRWh4xOCY

  • But they didn't only train on information the creators made freely available. They trained on copyrighted materials obtained illicitly.

    • I know we're not supposed to comment about downvotes, but the original comment was talking about "these companies", and none of the information indicating that they, or at the very least Meta, trained on terabytes of books downloaded from zlib and libgen and other torrent sites, is in dispute. So even if you believe that copyright should not exist, I don't see why this is not a valid dispute of the parents argument that they only trained on information creators made freely available.

  • > "Having machines consume large volumes of data posted on the Internet for the purpose of generating value for them without compensating the creators" is equally a description of Google.

    Quid pro quo. Those sites also received traffic from the audiences searching using Google. "Without compensation" really only became a thing when Google started adding the inlined cards which distilled the site's content thus obviating the need for a user to visit the aforementioned site.

    • I'm not sure quid pro quo even matters. A search engine is more like providing a taxi service. You're just taking people to a place.

      Now the AI summaries are a different story. One where there is no quid pro quo either. It's different when that taxi service will also offer the same service as that business. It's VERY different when that taxi service will walk into that business, take their services free of charge[0], and then transfer that to the taxi customer.

      [0] Scraping isn't going to offer ad revenues

      [Side note] In our analogy the little text below the link it more like the taxi service offering some advertising or some description of the business. Bit more gray here but I think the quid pro quo phrase applies here. Taxi does this to help customer find the right place to go, providing the business more customers. But the taxi isn't (usually) replacing the service itself.

    • Arguments like this never work out. There is no agreed upon compensation for being listed. If I didn’t want my site listed by Google and it was listed anyway, I may not think the traffic justifies my subjective “cost” of being listed. There’s also no legal protection against having my publicly accessible site and the title in its html from being shown (as there shouldn’t be).

We didn't seem to mind when Google was doing it back in 1999, or Lycos, Altavista, etc before them... why do we care about the LLM companies doing it now?

  • I find LLMs extremely useful but I think the difference is that they regurgitate the content (not verbatim) instead of a link to it. This is not unlike how a human might tell their friend about it.

    • Google has been regurgitating content right into search results since the very beginning, and they've been providing "synopsis" type of results for over a decade.

    • > This is not unlike how a human might tell their friend about it.

      Is there someone who has read the whole internet? Can we all be there friend?

      The entire basis of fair use is scale matters.

  • Because they have terms of service they have to adhere to. We need laws to be lawful.

I consumed large volumes of data posted on the internet for decades, which generated a lot of value for me, without compensating the creators.

The only difference is that I (presumably) have a soul.

> "Reading stuff freely posted on the internet" is also very different from a business having machines consume large volumes of data posted on the Internet for the purpose of generating value for them without compensating the creators.

The fact that value is being created is irrelevant. The fact that they are making profit is irrelevant. As is non compensation to creators. There isn't any law being broken. Is there?

Bottom line in real world terms there is no expectation of privacy with a freely open and unrestricted web site. Even if that website said 'you can use this for single use but not mass use' that in itself is not legally or practically enforceable.

Let's take the example of a Christmas light show. The idea might be (in the homeowners mind) that people, families, will drive by in their cars to enjoy the light show (either a single home or the entire street or most of it). They might think 'we don't want buses full of people who paid to ride the bus' coming down the street. Unfortunately there is no way to prevent that (without the city and laws getting involved) and there is nothing wrong with the fact that the people who provide the bus are making money bringing people to see the light show.

> "Reading stuff freely posted on the internet" is also very different from a business having machines consume large volumes of data

...not if you believe in the right of general-purpose computing. If they have the right to read the data, why don't they have a right to program a computer to do it for them?

I think we all agree that they're not the good guys here, but this reasoning in particular is troubling.