Comment by marssaxman

1 year ago

"Reading stuff freely posted on the internet" constitutes stealing now?

Seems like an excessively draconian interpretation of property rights.

"Reading stuff freely posted on the internet" is also very different from a business having machines consume large volumes of data posted on the Internet for the purpose of generating value for them without compensating the creators. I'm not making a value judgement one way or the other, but "reading stuff freely posted on the Internet" is an oversimplification.

  • Okay, but "stealing" is also an oversimplification, to the point of absurdity.

    It makes no sense to put stuff up on the internet where it can freely be downloaded by anyone at any time, by people who are then free to do whatever they like with it on their own hardware, then complain that people have downloaded that stuff and done what they liked with it on their own hardware.

    "Having machines consume large volumes of data posted on the Internet for the purpose of generating value for them without compensating the creators" is equally a description of Google.

    • They are not free to do whatever they like, there are tomes of laws across all countries governing what someone can and cannot do with your intellectual property. Just because we didn't have the foresight to add in a "if by chance in the future someone invents artificial intelligence, that's not fair use" is a shame, but doesn't make what these companies are doing ethical or morale.

      I don't disagree regarding Google, I also think they exploited others IP for their own gain. It was once symbiotic with webmasters, but when that stopped they broke that implied good faith contract. In a sense, their snippets and widgets using others IP and no longer providing traffic to the site was the warning shot for where we are now. We should have been modernising IP laws back then.

      10 replies →

    • What if that data isn’t publicly posted? For example, copilot regurgitating code from private repos, complete with comments.

    • Proper term for it is Computer Assisted Plagiarism, CAP for short. Also, I really hope that Google doesn't claim it created sites it crawl for search their engine.

    • Ok so if I publish under a license saying I don't allow for it to be used for AI do you believe they respect it? What word would you use to describe this violation? Go ahead throw up a robots.txt, throw up a license. You will be able to coax the "fair use" stochastic parrots to render it verbatim.

      Sam Altman and his ilk are exploiting the incredibly slow moving legal system to enrich themselves.

      1 reply →

    • It's not about the downloading of the data, it's about its use in training models, which is dubious from a copyright perspective.

    •   > where it can freely be downloaded by anyone at any time, by people who are then free to do whatever they like with it on their own hardware
      

      I think you have a strong misunderstanding of the law and the general expectation of others.

      I'd like to remind you that a lot of celebrities face legal issues for posting photos of themselves. Here's a recent example with Jennifer Lopez[0]. The reason these types of lawsuits are successful is because it is theft of labor. If you hire a professional photographer to take photos of your wedding then the contract is that the photographer is handing over ownership of the photos in exchange of payment. The only difference here is that the photo was taken before a contract was made. The celebrity owns the right to their body and image, but not to the photograph.

      Or think about Open Source Software. Just because it is posted on GitHub does not mean you are legally allowed to use it indiscriminately. GitHub has licenses and not all of them are unrestricted. In fact, a repo without a license does not mean unfettered usage. The default is that the repo owner has the copyright[1].

        > You're under no obligation to choose a license. However, without a license, the default copyright laws apply, meaning that you retain all rights to your source code and no one may reproduce, distribute, or create derivative works from your work.
      

      A big part of what will make a lawsuit successful or not is if the owner has been deprived of compensation. As in, if you make money off of someone else's work. That's why this has been the key issue in all these AI lawsuits. Where the question is about if the work is transformative or not. All of this is in new legal territory because the laws were not written with this usage in mind. The transformative stuff is because you need to allow for parody or referencing. You don't want a situation where, say... someone including a video of what the president has said to discuss what was said[2]. But this situation is much closer to "Joe stole a book, learned from that book, and made a lot of money through the knowledge that they obtained from this book AND would not have been able to do without the book's help." Just, it's usually easier to go after the theft part of that situation. It's definitely a messy space.

      But basically, just because a piece of art exists on public property does not mean you have the right to do whatever you want with it.

        >  is equally a description of Google.
      

      Yes and no. The AI summaries? Yeah. The search engine and linking? No. The latter is a mutually beneficial service. It's one thing to own a taxi service and it is another to offer a taxi service that will walk into a starbucks take a random drink off the counter and deliver it to you. I'm not sure why this is difficult to understand.

      [0] https://www.bbc.com/news/articles/cx2qqew643go

      [1] https://docs.github.com/en/repositories/managing-your-reposi...

      [2] https://www.youtube.com/watch?v=tUnRWh4xOCY

    • But they didn't only train on information the creators made freely available. They trained on copyrighted materials obtained illicitly.

      1 reply →

    • > "Having machines consume large volumes of data posted on the Internet for the purpose of generating value for them without compensating the creators" is equally a description of Google.

      Quid pro quo. Those sites also received traffic from the audiences searching using Google. "Without compensation" really only became a thing when Google started adding the inlined cards which distilled the site's content thus obviating the need for a user to visit the aforementioned site.

      2 replies →

  • We didn't seem to mind when Google was doing it back in 1999, or Lycos, Altavista, etc before them... why do we care about the LLM companies doing it now?

    • I find LLMs extremely useful but I think the difference is that they regurgitate the content (not verbatim) instead of a link to it. This is not unlike how a human might tell their friend about it.

      2 replies →

    • Because they have terms of service they have to adhere to. We need laws to be lawful.

  • I consumed large volumes of data posted on the internet for decades, which generated a lot of value for me, without compensating the creators.

    The only difference is that I (presumably) have a soul.

  • > "Reading stuff freely posted on the internet" is also very different from a business having machines consume large volumes of data posted on the Internet for the purpose of generating value for them without compensating the creators.

    The fact that value is being created is irrelevant. The fact that they are making profit is irrelevant. As is non compensation to creators. There isn't any law being broken. Is there?

    Bottom line in real world terms there is no expectation of privacy with a freely open and unrestricted web site. Even if that website said 'you can use this for single use but not mass use' that in itself is not legally or practically enforceable.

    Let's take the example of a Christmas light show. The idea might be (in the homeowners mind) that people, families, will drive by in their cars to enjoy the light show (either a single home or the entire street or most of it). They might think 'we don't want buses full of people who paid to ride the bus' coming down the street. Unfortunately there is no way to prevent that (without the city and laws getting involved) and there is nothing wrong with the fact that the people who provide the bus are making money bringing people to see the light show.

  • > "Reading stuff freely posted on the internet" is also very different from a business having machines consume large volumes of data

    ...not if you believe in the right of general-purpose computing. If they have the right to read the data, why don't they have a right to program a computer to do it for them?

    I think we all agree that they're not the good guys here, but this reasoning in particular is troubling.

I'm not talking about that, I'm taking about downloading gigabytes of books, and movies and who knows what data (since it's not disclosed) without paying. Those are not freely posted on the internet. Well, not legally anyways.

This is a quintessential bad faith comment.

The reference to terabytes of stolen data refers to copyrighted material. I think you know this but chose to frame it as "stuff freely posted on the internet" in order to mislead and strawman the other comment.

  • I meant it exactly as I said it. I do not agree that any theft occurred, either in law or in spirit, and I believe that reinterpretation of intellectual-property law in order to make it a crime would cause significant harm, greatly outweighing the benefits, as has been the case with every other expansion of intellectual property law I have seen.

Faithfully reproducing something you've previously read while passing it off as your own original work is a violation of the most basic tenets of intellectual property rights.

Forgot the 82TB of torrented books Meta has been using for training? I mean, yeah, it’s Meta. No surprise. But I won’t believe for one second that the other players didn’t do a similar thing. They just haven’t been caught yet.

so I can take a screenshot from a movie trailer on YouTube and sell posters of it now? I thought copyright still applied to the poor.

What "reading"?

  • The same reading search engine crawlers have been doing since time immemorial.

    • No one gave them permission to access their webservers back then either. Before it's cited that there is precedent in law, that is in the US. No such precedent exists in my country, and our laws suggest that unauthorized access regardless of "gates up or down" would constitute trespassing. There are also no protections for scrapers coming out of prior lawsuits, and copying copyrighted material is of course illegal.

      Which is just to point out that the world wide web is not its own jurisdiction, and I believe AI companies are going to be finding that an ongoing problem. Unlike search, there is no symbiosis here, so there is an incentive to sue. The original IP holders do not benefit in any way. Search was different in that way.

    • Search engines never claimed that their content was orignal, and redirect to the original author (which gets proper retribution)

"Reading stuff freely posted on the internet" that has copyrights to be used in your generative AI service is stealing, is a pretty basic interpretation of property rights.

As long as people are being prosecuted for piracy or having their livelihoods compromised for including a 16 second clip of a song, yes.