Comment by Aurornis
7 days ago
They cannot do those things. The reason they’re scanning the books is because it’s not possible to legally obtain or transfer their digital copies. They have to do their own scans and keep them in house.
7 days ago
They cannot do those things. The reason they’re scanning the books is because it’s not possible to legally obtain or transfer their digital copies. They have to do their own scans and keep them in house.
First, to be clear, I work for Amazon, and I don't work on this or related efforts, I am a security engineer. I also won't talk about or answer questions at work, and my comments are more generally about the practice (all of the major organizations building AI are doing destructive book scanning). These are my opinions, and do not reflect my employers (past or present).
You might be right that they can't do that now. The simple path forward is to have these companies simply make a public commitment to publish the data when the copyright expires. I also think there is a space to be carved out, probably through regulation, to ensure that there is a clear path for these scans to enter the public domain, at the very least.
This is not just important for these specific books, I have written on other platforms and in other spaces about the importance of media companies and those who benefit from strong copyright laws to protect and generate profits and revenues to repay the public for the cost of that enforcement over time by ensuring that at the appropriate time, those works fully enter the public domain. That could mean a restructuring of the Library of Congress in the United States to become a modern Library of Alexandria to host data, and shifting to a registered copyright model where to gain the protections of the court, you need to upload/submit your copyrighted works for storage and eventual release. I doubt it would ever happen there because of the amount of money invested in tying up IP in the United States, but perhaps a more amenable location like the EU could help with that.
However it might work, part of the promise of the Internet was that information would be liberated, but we see every day how much information gets sent down the memory hole when businesses, sites or services shut down, or how regularly companies abuse IP related regulations to attempt to strangle competition. It would be expensive now, but it would create an incredibly valuable legacy of information for the future, and it will only get more expensive to build such a thing as time goes on.
Great comment. Those who try to excuse book burning cannot be doing it in good faith, this is the acid test.
If AMZN, et al. were burning books in good faith, they would be open about it, point to the party responsible for it, and provide a list of the burned books to allow the community to organize, scavenge and scan the endangered publications. Keeping everything secret is a proof of maliciousness.
Preserving books and providing easy access to the information in them is of utmost importance, this is certainly understood by the public and private entities who could do something positive about it - their actions in the opposite direction should be a wake up call, they're on the wrong side of this issue.
Libraries I volunteered for would coordinate throwing away books late night under dark right before the dumpster was picked up and emptied.
Because people are irrational when it comes to this topic. They say stuff like “book burning” when they see books being destroyed. Doing it in relative secret kept the crazies away.
If you are passionate about this topic, lobby for more funding to store archives in the public good. Expecting private parties to do it for free because it gives you the ick to see books destroyed is not useful.
More books get destroyed each year during estate cleanouts than any AI companies could ever hope to accomplish. Most books donated to goodwill or other thrift shops go straight to the dumpster and might not even have a set of human eyes put on them at all. These places act as sin eaters for folks to leave their trash with.
Copyright does expire, for older books.
It would make more sense for government to accept digital copies for any book, and share the out-of-copyright ones. In the United States, we have a Library of Congress who could do it if a law were passed with funding.
What they could do, if they were truly focused on not taking the most destructive path forward, would be to cooperate on a single scanning company that could pool the resulting text for the others on attractive terms, to ensure that once a book was scanned all the others wouldn't need to scan and destroy the same title. That would substantially limit the damage.
Of course, because each company is happy to burn it all down to beat their competitors and being first to any content is most important, this will never happen.
Somehow people are still failing to understand that transferring the scan from a trust, scanning company, or other external party would constitute a transfer of a copy between distinct legal entities that breaches copyright law in an actionable way.
Doing the scan within their own organization does not because it's within a single legal entity and there is no copyright issue involved.
I'm not failing to understand that, I just don't think it's an insurmountable barrier.
You could overcome it through various means, including helping to set up a clearinghouse for bulk rights, implementing a data clean room approach that models could train on in situ, lobbying Congress for copyright law changes especially around orphan works, using Section 108 of the Copyright Act to set up a specific preservation vehicle like the HathiTrust, and other options.
I find it ridiculous that so many in this thread are acting as though AI companies are simply powerless to do anything but buy up and destroy these rare books.
3 replies →