Comment by ThrowawayR2
1 year ago
You got suckered by the clickbait. Destructive scanning (https://en.wikipedia.org/wiki/Book_scanning#Destructive_scan...) isn't unusual for books that are common enough that an individual volume is of no particular value.
I didn't get suckered by anything. I'm aware of the practice. I find it objectionable. That they did this is just another thing on the growing list of objectionable things that genAI companies seem to enjoy doing.
To be honest, I probably wouldn't have even commented on it if it were the only bad thing these companies do.
It was only legal because they did it this way.
> Ultimately, Judge William Alsup ruled that this destructive scanning operation qualified as fair use—but only because Anthropic had legally purchased the books first, destroyed each print copy after scanning, and kept the digital files internally rather than distributing them. The judge compared the process to "conserv[ing] space" through format conversion and found it transformative.
Very laws that the publishing industry has lobbied so heavily to make so strict are the reasons for this behavior.
If you believe that destroying books is bad, your issue is with copyright law, not the AI companies. The AI companies are just following copyright law -- they are allowed to move data from one format to another (thereby destroying the original), but not copy it.
Not everything objectionable or unethical should or could necessarily be outlawed. "It's not illegal" is not really an argument or justification for anything.
1 reply →
> If you believe that destroying books is bad, your issue is with copyright law, not the AI companies
No, my issue is with the companies that do this. The law doesn't enter into it. Just because a thing is legal doesn't mean it's OK.
Specifically his issue is with First Sale doctrine. If you own it you can destroy it and its none of anyone else's business.
1 reply →
I very much have a problem with both of these things.
I mean, they could have gotten e-book versions of the books, or even preprint PDFs.
In an era where people are starting to calculate the environmental impact of the jobs they run on the cloud and start to optimize it, adding that much load on recycling system is not a wise choice, but only a selfish one.
I strongly suspect that dealing with ebooks on this scale might actually be even more onerous than the physical volumes.
The physical stuff is straightforward. Buy books from bulk sellers, rip off everything and put them into off-the-self rigs for digitization. It's straightforward, directly scalable, can use any book, and your main issue is format shifting, which anthropic successfully argued here. No DRM, you buy exactly the books you need, and every book is processed exactly the same way.
If you try to buy ebooks, you get wrapped up in onerous licensing terms about copying, and how you're able to use them, how long you're able to access them, and so on. Many books won't even be available (or can only be licensed alongside a bunch of others) and you have to deal with DRM you can't strip without creating additional copyright issues.
We've somehow created a world where physical objects are more free than bits.
No, they probably couldn't have. eBooks are notoriously DRMed and the DMCA makes it illegal to circumvent an effective copy protection mechanism even if you otherwise have legal access to work. Furthermore, first sale doctrine doesn't apply to any digital files and they can't be obtained legally in bulk.
I'm sure they would have loved to save the hassle and expense of disassembling physical books. Presumably something legal related or cost related prevented them from going that route.
Yes, they did it as a workaround for copyright. TFA explains that aspect.
1 reply →