← Back to context

Comment by whatever1

7 days ago

This is what the copyright laws dictate no ?

Bias disclaimer: Amazon is my current employer, but I don't work on AI or anything else mentioned in the article.

Yes, this is a result of copyright laws. The other commenters are wrong/uninformed.

If it was up to the companies training LLMs, they wouldn't destroy the books: It's a waste of company resources, it's needlessly destructive/evil, it generates bad PR, etc etc. There are essentially zero advantages, other than it is what is required under US copyright law (or at least, it is what their highly paid lawyers believe is required under US copyright law).

  • Isn’t it being destroyed because it makes the scanning process easier?

    • Perhaps partially? I assume that cutting the pages out of the spin makes them easier to scan at least partially. That said I have no insider knowledge of this type of operation so I don't know how much easier that actually makes it.

      But ultimately, it's a moot point, because the legal requirement means the books must end up destroyed. Even if the people at Amazon wanted to scan the books in a way that required no destruction at all, it's not currently (legally) possible for them to do so, so they might as well take the easy way out today.

      2 replies →

    • No.

      A judge a while ago decided that as long as the physical copy is destroyed, and "transformed" into an electronic copy, you can do the upload. But if you preserve the physical copy after scanning it, you are in violation of copyright because you "copied" the book.

      That's literally the only reason they are trashing them. It's a legal requirement.

      4 replies →

Not to my understanding. To begin with, it's far from given that "rare" books are all covered by copyright. But if they are, it's at best murky: whether you destroy the original doesn't really have anything to do with what you're doing with scanned contents. The scanned contents themselves may be inherently a copyright issue, regardless of destroying the original. The actual trained model has separate arguments more in its favor, so if no scanned contents exist - IE the data is read once for training and not stored or saved, they have a better argument. But in that case the destruction is totally disconnected from copyright, as they'd be totally okay to rescan the material.

Those laws are lobbied for by large corporations, these are not just laws that exist outside of that context. They can also be changed, or Amazon could just incur the fines.

Large corporations will move fast and break things when it’s convenient; they don’t care much about the law - just about profit.

i would assume that would come into play if they were uploading scans of the books? there must be some gray area where training like this doesnt apply to that. or ya know, just do it and face the consequences later because you already have the data and know that ai obsessed government will just shrug their shoulders.

No. A copy is still a copy even if you destroy the original.

  • Until about a year ago this would have been a reasonable and respectable argument, but at least in California you are arguing against current legal precedent: https://fingfx.thomsonreuters.com/gfx/legaldocs/jnvwbgqlzpw/...

    • That's a bit of a wild interpretation of copyright law.

      I mean, Anthropic isn't going to fight it because it lets them do the thing they want to do, so I can see how this never gets beyond the court that allows them to do the thing they want to do.

      But would this argument would have flown in the past?

      It wasn't even attempted in Sony v Universal. Or any copyright suit up until this point. That doesn't smell funny to you?

      1 reply →