I really have trouble getting worked up about this. "Rare books" is thrown around regularly but my gut feeling is that's not the case. These are used (often? always?) books and while I'm sure there is waste, in general they just want 1 of every book.
While I wish there was a repository of every book that was already digitized (it pains me this is the best solution), there isn't one and so I think this is not a real problem.
It'd be a different story if they had furnaces that ran only on rare books that they had to continually feed books to but that's not what's happening here. And that 1 destroyed copy will "live on" in a way that it otherwise might not.
It'll live on if they publish those scans or contribute them to a national archives or something. Proprietary data has a habit of being lost over time though.
They are literally not allowed to publish those scans.
Half the reason the books get trashed in this process is because the first sale doctrine keeps copyright from strangling all the freedom in this narrow area.
It seems a little short-sighted to destroy an artifact to get the text.
Perhaps the genetic material that remains in books from the people who handled them, or the pollen from plants in the environment that the book existed in will have value in the future, but we won't know what we lost because some people foolishly destroyed it in a bizarre quest to make AGI that the creators argue could potentially destroy humanity.
The more and more I read about these kinds of people the more I'm starting to realize that they're in the "here for a good time not a long time" group of people and those are the last people you want making long-term decisions.
We run a little library in front of our house. It's amazing how many people dump boxes of old books off in front of it hoping that they will find their way back into someone's collection.
The sad truth is the vast, vast majority of printed literature is neither interesting nor useful. People are not dumping off stacks of Umberto Eco. We frankly have to toss a lot of awful cookbooks, self-help books, trashy mass-market "novels", and sketchy religious works. As it is, even the stuff that makes it to the library is not very impressive.
Imagine if we had bad cookbooks, cheap popular novels and stories, tracts on diet and self help from old Rome, or old eras in China, or the equivalent from ages before that.
It's all interesting for something, even if it's just a meta analysis of culture during a certain period or what kind of trashy romance novels were popular in 198X. At least in my view.
Do we know if they actually buy one copy of each book here? Or are they indiscriminately buying books in bulk and just processing all of them?
The latter seems inefficient, so my first assumption would be that they would avoid that. But while I suspect they'd check if they know the book before scanning, I could imagine them not caring that much before buying them and just focus on volume.
Books are not original manuscripts. Even in low volume cases, they are usually printed hundreds of times. (And usually low volume works aren't all that great...hence the low demand.)
That is a fraction of a percent for books that at some level weren't all that wanted.
The Embassy of the Free Mind (https://www.embassyofthefreemind.com) is a rare book library in Amsterdam that is scanning books the old fashioned way… leading to https://SourceLibrary.org — a collection of over 5,000 books from the renaissance that have never been translated before. Consider donating, if this is a topic you care about!
The length of this article and the unnecessary visualizations felt like a waste after getting to the end and discovering they won’t reveal anything about the books that were scanned.
The closest they got was admitting that the rare books weren’t anything that someone might care about in the sense that people assume when we hear “rare books”
> As the bookseller who sold them told me, there are not many people in the world who would care about them in the same way people might care about the first edition of Oliver Twist, but that doesn’t mean they’re not valuable.
Okay? But then why exactly where they considered valuable enough to warrant an entire article about them going to a book scanning facility? Without revealing anything about these books I have no idea if they were classic literary works that were underappreciated, or if this was some old guide about How to Use Microsoft Office 97.
There was a more balanced take on Twitter (which I’m unable to find again, because Twitter) from a book seller who said it was more of the latter type: Books that were rare because they were no longer in demand and most everyone had thrown their copies away. Some parts of the media are doing backflips to try to imply that these are cherished literary classics being fed into the shredder to deprive humanity of something valuable, but the book seller seemed happy to be making sales for useless old books that no human was interested in buying.
As if the AI industry is only shredding "crappy" books.
I recently bought a rare book. It was a boat design book written by a famous yacht designer in the 1940s, but it's been out of print for decades and I had to pay $150 for a "fair" copy with missing dust jacket.
I have an interest in older technology and the old ways of doing things (for instance, how do you lubricate the mast of a gaff rigged boat so the gaff jaws don't jam?) I've often found myself reading very old books that have been out of print for a century.
Sometimes I read those books at libraries and I've been the only person to check them out in years (I started doing this back when they still stamped the return date on a card so you could see when it was checked out). Now most of those books have been disposed of by libraries due to yield management software and I've ended up with some of them in my personal collection, but people like me can only save a tiny sliver when most of them are being bought in mass by Sam Altman. When they're gone the knowledge in them is also gone.
We are burning the library of Alexandria and the HN consensus is "those books probably weren't saving anyway."
> As if the AI industry is only shredding "crappy" books.
Do you, or anyone else, have any source suggesting that they’re buying highly valuable rare books and shredding them?
The kind of $150 rare book that you had to buy from a specialty collector who graded it is in a completely different category. You’re thinking of “rare books” in the historically rare, valuable, and collectible category.
The book sellers shipping off orders of 1000s of books at a time to these facilities are calling the books “rare” because they may only have 1 copy, not because it’s a collectible with a high price tag.
It would be nice if your local library hadn't sold their copy but I think your beef is with the libraries, or perhaps with the politicians who failed to fund them in that case.
Actually I expect that a copy of most books probably is still available in a copyright library but it might not be easy to access.
Weeding (the library term of art for selecting works for deacquisition) is a very import part of collection management. It's actually regionally and nationally coordinated so the interlibary loan network does not throw away the last copies.
Once you start to realize just many books got thrown out annually before “AI” you have less sympathy. Libraries are one of the biggest contributors to this because there is simply too many books that nobody cares about.
I suspect folks are over weighting how much knowledge is being destroyed in these books. If someone actually quantified it that would be amazing but as someone who started going to used book sales at a very young age I just have no sympathy. Most books are worthless. I don’t mean that from a text perspective either.
> or if this was some old guide about How to Use Microsoft Office 97.
Materials like that may be very valuable to software / tech / HCI archeologists soon.
I can't get the tools or local know-how to straighten my scythe blade in a country where every cottage had a scythe less than 100 years ago, with the last scythe-native generation rapidly dying out.
Fast forward a civilizational collapse and that M$ Office 97 for Dummies might be as groundbreaking as a Guide to Using Roman Concrete
Anyone who is into books eventually finds out that we're permanently losing them all the time. Like they're thrown out, and lost forever. University libraries throwing out huge collections of out-of-print material to make room for new books, study spaces, even cafés. Municipal libraries turning over their collections. Books that never made it to libraries going out of print, tossed in the trash after yard sales.
The AI companies digesting this stuff is a net win for humanity. And I'm not a fanboy! Ideally they'd upload them to Anna's archive too, but even if they keep it private forever, at least these books live on in some way in the model weights. Thats better than a landfill.
It would be really great if the companies who are doing this would commit to placing the scanned files into a public trust that would coordinate with organizations like the Gutenberg project to ensure that the scanned materials enter the public domain on schedule. Publishing encrypted archives with the keys in escrow would be a good first step.
IMO that would go a long way to resolve any concerns about losing books. I still don't like the idea of extremely hard to find or last prints being actually destroyed for this, but it certainly makes it more palatable.
I would feel much better about this process if they were uploaded and if it were framed as a knowledge preservation project. This would only slightly increase the cost of the project, but have a huge impact on its perception and its net positive impact.
Of course, actually benefiting humanity is only a minor, indirect concern for investors.
If or when they go bankrupt or reach agi, they will just delete them. I hope Anna's archive already has them anyways. Apparently openai's newest unreleased model, gpt 6, is capable of continuous training at Inference time, like a person is. That might be enough to delete them
That is the problem: they are digesting it. They are not creating a new kind of library, where you could say, show me the text of "How to Fix Your Ice Cream Problems". (An actual book I own) It is not being done for our future reference.
> Ideally they'd upload them to Anna's archive too
Not only can they not do that, they must scan physical copies because they are forbidden from using digital pirated copies from sources like this.
Anthropic had a big settlement because they were caught using downloaded digital copies. As a response they’ve ramped up their book scanning and others have followed.
This is the right answer but the reason they can’t upload the scans is copyright law as demonstrated by Google having to settle with the publishers and allow them to remove their books and limit free access to 20% of text. The AI companies are essentially compressing the information in a huge swath of books that would otherwise be headed to landfill and making them 1000x more accessible. This is unquestionably one of those instances where capitalism is taking money from rich investors and benefiting the 99%.
This feels like a manufactured controversy. What difference does it make to me what someone does with a book after they buy it? It's effectively unavailable to me regardless of what they do. If people are really concerned about these "rare" books, they should lobby the copyright owners to release them online or print more copies.
> What difference does it make to me what someone does with a book after they buy it?
I think this is a bit of a myopic take. It's like saying it's my property I can do what I want, yet there exists designated historic homes or neighborhoods that are deemed to have cultural value where modifications do in fact need to be approved.
I'm with you on the "rare" part. If people are thinking about 70+ year old documents or ancient manuscripts, I doubt that's what AI is being trained on and is being destroyed, but it's reasonable that people find _that_ idea distasteful.
You can say it's manufactured but if these companies ignore this criticism, it's just another way AI companies are committed to losing the public.
> It's like saying it's my property I can do what I want, yet there exists designated historic homes or neighborhoods that are deemed to have cultural value where modifications do in fact need to be approved.
Yes, and when you buy them with that designation, you know what you're getting into. The problem occurs when you already own it, and some group is trying to get it labeled as historic, which will add to your burden and limit what you can do with it.
When I was in a small town, this was actually weaponized. A hotel owner was trying to get another hotel categorized as "historic" and had rallied a lot of people behind his cause. He had a case - the hotel did have some claim to being the "first" in some category or other. But really, he was doing it because it was a competitor. The "historic" hotel owner had to spend a lot of money to fight the cause, because being labeled historic would prevent him from performing various upgrades, making the hotel less attractive to customers (he was already not getting many customers).
Replace bison with some random animal no one's ever heard about, and photos with literal clones of the animal, and you've got a much better metaphor. Not to forget that in this world, brand new animals are created every day.
This is a problem whether it's AI companies buying them or three random dudes. The real solution is to get the copyright owners to keep distributing copies, or to change copyright law.
I think the term "rare" might be a bit loaded. DOES it mean "only a handful in existence" or more like "1000 copies?" And at any rate, at least by scanning the book they're theoretically making its
contents available to the public. What are the other people who own these books doing besides having them sit on a shelf?
> We’re not revealing the titles of the books included in the shipment we tracked, but they are rare, meaning there are not many copies of them in circulation. Sometimes that’s because not many copies of them were ever printed, and sometimes because they are in a foreign language not many people speak.
> The seller agreed to put an Apple AirTag provided by 404 Media in one of the books included in this order so we could see where the book was going. And by extension, which company, AI or otherwise, was behind this massive order.
It would seem the redaction of the rare titles is a way to avoid de-anonymization and subsequent harm to the business of the seller who agreed to place a tracker in one of the books. That being said, maybe they could have chosen a better methodology which would have allowed the disclosure of the title, although ultimately I’m not sure the title matters too much outside of their claim they were “rare”.
The issue is if rare is simply "not many in circulation", that's very very different from rare being "it is a unique artifact". The former kind of rarity doesn't really present an existential end, whereas the latter does.
And even then, the "rare" qualifier is not needed here. What Amazon/LLM companies are doing is amoral. This is intellectual piracy (not in the "copyright infringement" sense) at the highest level. Stealing and centralizing the accumulation of human knowledge to eventually rob us all and put all power into the hands in the hands of a handful of people, who are not benevolent.
It doesn't really matter. The article suggests that they are selecting and tracking books by ISBN, which means: books that have been published or at least reprinted in the last 50 years or so. And they are trying to get as much of that set as they can, regardless of quantity, quality, or any other consideration. Which means that older books printed before ISBNs became common may be relatively safe, at least from Amazon. And some books just don't come up on the used market very often.
(small historical irony: when Amazon first started selling books, they used the Books in Print database, which included a lot of books not actually in print.)
Rare as in your grandfather's John Deere manual from 1982, not rare as in a test print run of The Great Gatsby. The number of books ingested by these AI companies is a drop in the bucket compared to old books destroyed every year through normal means.
> The number of books ingested by these AI companies is a drop in the bucket compared to old books destroyed every year through normal means.
Citations? Also what exactly are these 'normal means'? As a bibliophile who loves scouring used book stores for out-of-print titles this is a topic I'm very interested in.
EPA estimates that hundreds of thousands of tons of books are landfilled or recycled every year. That's just the United States. It seems likely that the worldwide figure is in the millions of tons.
Normal means is throwing it in the recycle bin. Especially for stuff like a 1982 John Deere manual. I've never donated an old appliance's manual to the library. Have you?
Amongst my friends, I'm one of the rare folks who donates books to the library. Most people just trash them. And I know the library only wants them to try to sell them in their book sales (or online) so they can get money. Almost nothing one donates to a library actually ends up on the library shelves.
The scariest thing for me about companies not caring even minimally about conservation is that when AI gets more powerful than humanity (which is clearly a when not an if, even if there's a lot of uncertainty and differences in opinion about the time horizon here), I want to hope that AI will care more about conservation of human people.
So far from how I see how powerful organizations work, I'm not as certain as I would like to be.
> AI gets more powerful than humanity (which is clearly a when not an if, even if there's a lot of uncertainty and differences in opinion about the time horizon here)
People say this all the time, but so far nothing has convinced me it's true.
LLM development has more or less plateaued, and the current boundaries are very real - energy, resources, capital.
At this point we're talking about marginal improvements against the same asymptotes of all technological innovations.
So, if a corporation can suck up entire books to teach their machines how to think using the information from those books, can we (all humans) join a single corporation that provides all books to its employees? Just need one copy of each and we'll make that copy digitally available for our employees so they can learn from and utilize the knowledge from the books. It's not copyright infringement, they're employees.
Bias disclaimer: Amazon is my current employer, but I don't work on AI or anything else mentioned in the article.
Yes, this is a result of copyright laws. The other commenters are wrong/uninformed.
If it was up to the companies training LLMs, they wouldn't destroy the books: It's a waste of company resources, it's needlessly destructive/evil, it generates bad PR, etc etc. There are essentially zero advantages, other than it is what is required under US copyright law (or at least, it is what their highly paid lawyers believe is required under US copyright law).
Not to my understanding. To begin with, it's far from given that "rare" books are all covered by copyright. But if they are, it's at best murky: whether you destroy the original doesn't really have anything to do with what you're doing with scanned contents. The scanned contents themselves may be inherently a copyright issue, regardless of destroying the original. The actual trained model has separate arguments more in its favor, so if no scanned contents exist - IE the data is read once for training and not stored or saved, they have a better argument. But in that case the destruction is totally disconnected from copyright, as they'd be totally okay to rescan the material.
Those laws are lobbied for by large corporations, these are not just laws that exist outside of that context. They can also be changed, or Amazon could just incur the fines.
Large corporations will move fast and break things when it’s convenient; they don’t care much about the law - just about profit.
i would assume that would come into play if they were uploading scans of the books? there must be some gray area where training like this doesnt apply to that. or ya know, just do it and face the consequences later because you already have the data and know that ai obsessed government will just shrug their shoulders.
I support physical media and doing whatever you want with said physical media. If you want to overpay for a copy of Sharepoint 2007 For Dummies and destroy it, knock yourself out. Just because media has been printed and is "rare" doesn't mean it has any practical value.
You could easily write an article about how Goodwill and the Salvation Army dump millions of "rare" books that no one would ever buy for $0.50. And thrift stores will dump multiple copies of said books. AI companies would only ever need one!
The only thing I disagree with is AI companies feeling entitled to freely use every piece of copywritten work without restriction, but that applies just as much to web scraping, pirated books, etc.
Why is the textual data bound up in these paper books worth the trouble?
These LLM training runs have already ingested essentially the whole public internet. What marginal value is to be gained from scanning and destroying obscure books?
For me the biggest functional issues with LLMs don't seem to have any connection with "I wish they had read this obscure community cookbook from 1946". Is that going to get Claude to stop saying "honestly" to me? Is it going to get LLMs to stop making up sources that don't exist? What is in it for Amazon or any LLM company to chase more obscure data like this.
Frontier researchers have found that dumping more and more data into training is effective at improving LLM capabilities. Nobody has a gears-level understanding of how training on some particular kind of data leads to some particular behaviors, so they generally take the attitude that more is better.
I have to agree; even if they are getting regular 20th century out-of-print books, is that going to add a significant percentage to their training data?
I can only think of it being a 'low-background steel' situation where they want to locate original, non-digitized text for validation or knowledge bases.
I'm not an expert in LLM training, but I think we can all agree that the writing on the Internet is generally very low-quality and surface level compared to the depth of books. Most books in the past were even edited by a separate person from the writer!
Humans manage to be pretty intelligent with only reading perhaps a few thousand books in their lifetime. It seems unlikely that AGI will appear but only once it has read that last out of print 1983 book on knitting patterns.
Granted that that is indeed the case, surely the quantity of high-quality writing available online (including things like Project Gutenberg) still dwarfs the quantity remaining in obscure unscanned books.
We really need to be thinking about the long-term implications of what Amazon, Anthropic, OpenAI, and others are doing here.
Because it's not just one company doing this, it's dozens if not hundreds of them, all with huge budgets and all competing over the same dwindling supply of older books. What those talking about library disposals and similar measures are missing is the sheer scale that destructive book scanning for AI ingestion is operating with. It's unprecedented.
It's not impossible that in a few years the number of older physical books for sale anywhere will plummet to almost nothing and that in many cases that'll include the last copies anywhere of particular titles. Second-hand book stores will close, and a ton of niche knowledge might be lost forever. It'll also all have been done in relative secret, without any transparency or public record kept of what was lost.
Worse, it's something that can only be done once. If our societies don't do something about this now there isn't going to be a chance for a do-over. Once a physical book is destroyed it's gone forever, and we'll be lucky if there's a digital copy left. But even then digital copies don't provide the same level of forensic verifiability that physical copies do. Future researchers looking for material that might've been contained in books like these will be out of luck.
If our societies do nothing to stop this, we'll all be poorer for it.
I don't understand why Amazon still bothers to blow money on their in-house AI effort (outside AWS Bedrock). What have all the scraping, scanning and training led to? Nobody is using their models.
A lot of the rare books seem to be books that are most relevant to specific professions within specific countries. I would expect university or national libraries in those countries to have copies of those books, but they often do not digitize them because controlled digital lending is seen as a legal grey area at best. So these ai companies are doing digitization for them and not publishing it, duplicating work for these libraries once they do decide to digitize.
Many other rare books are just rephrasings of other books on a given topic and this can be done with synthetic data generation now. Same goes for websites
Honestly this feels like a US exclusive problem. Like shipping from overseas would be way too expensive for the cheap books they want. The Chinese just use Anna's archive instead
* They're only buying a single copy of that book as they only need one to scan
* If the book was public domain (or should be), then there should be an effort to "democratize" that data into a public commons of intellectual property?
Would that an enhancement to the Library of Congress or such?
Why do you think for even a moment that they would buy only one copy, rather than every available copy of something rare and difficult to get? Anyone doing this stands to gain from destroying the only copies of something that has scarcity, to stop competitors from ever getting access to it. That's a serious concern.
Why would they do that? What is this 'should' to a company that would do this in the first place? Are you not suggesting the opposite of what they aim to produce?
That seems really speculative. I think you're assuming the maximum possible evil intent for no reason, mostly because you hate AI and anyone involved in it.
While we can quibble (and are) over the details of this particular occurrence, I'm beginning to take the opinion that used/rare booksellers need to implement some form of "KYC" to help guard against the epistemic threats posed by this practice.
The booksellers are probably so ecstatic to get some value out of books otherwise destined for the landfill that they wouldn't lift a finger for "epistemic threats".
Years ago I saw an article on this. For Google Books, they had two processes.
The first was destructive. This was for mainstream books currently being published so they had no value. It's (I believe) where you cut off the spine and scan the pages.
For rarer books, there was a non-destructive process. Basically the book was opened to each page and scanned. This was slower but didn't destroy the book.
I don't understand why these companies haven't just licensed the scans Google has already done. Why is each company doing this rather than just scanning the books once and sharing the scans?
> I don't understand why these companies haven't just licensed the scans Google has already done.
Because for the vast majority of the books, Google doesn't have the legal right to license those scans. It would be legal if the books were out of copyright, but despite the connotation of "rare books" those generally aren't the books we're talking about. Further, in the cases where Google didn't destroy the original of the book they scanned, their scanned copy may be considered infringing under the new standard, so Google doesn't want call undue attention to what they have.
6 Feb 2021
At the Internet Archive, this is how we digitize a book.
We never destroy a book by cutting off its binding. Instead, we digitize it the hard way--one page at a time.
Wonder the additional cost to add automated page turning. More than Bezos can afford, impoverished chap he is.
If I remember correctly, it wasn't a distinction between common and rare books.
It was a distinction between library books and others. They partnered with libraries, and obviously, libraries didn't want their books destroyed, so they devised a non-destructive book scanner. Some of the books they scanned were indeed rare, but rare or not, you don't destroy books you borrowed from a library!
According to the article, the scanning facility has been there for a while, so it may well be that they have been destructively scanning books for a while, even before AI. It makes sense for the major book retailer to have digital copies of every book so that it can begin selling those as soon as it can negotiate the rights (if there are any).
In any case, all the indignation about destructive digitization misses the point that rare books takings space in a warehouse for years without being bought will eventually be destroyed anyway.
Something I've been thinking about is that a physical book is a physical artifact, and it's age and authenticity can be verified by treating it as such.
In contrast a digital book has no such marks. It's impossible to tell if it was redacted to hide inconvenient passages, or even completely rewritten or fabricated (which can now easily be done at scale with AI). If the physical book is a primary source then the digital version must be considered a secondary source: A potentially biased retelling.
For this reason, a digital book can never be a perfect substitute for the physical book it was created from.
I'm less concerned with how these rare books didn't rot in large lots of unused books and more concerned with whether or not this helps preserve the content longer.
I seriously doubt any of the companies doing this will preserve the scans for very long. It costs money to store stuff, and none of these companies are have any concern for anyone who isn't them, so they're not going to spend the money or lift a finger unless they work out a way to make it profitable.
archive.org estimates it costs $2/GB to store data in perpetuity (well, at least until decades of storage scaling trends stop of course). That's probably less than it costs to acquire and scan the data in, I'd be surprised if they just tossed it out at the end when it's so cheap to store. The same thing happened at the healthcare data lakes I worked on where the trend switched from the usual "how long are we legally required to store this information" to "how long can we legally hold this information".
I don't expect an amazon.com/checkoutrarebooks page but I also don't really think the legality of sharing the content has been much of a concern for these companies either. Ironically, one of the few things Meta got in trouble about with building their AI was helping share the training data.
I look forward to the day that AI companies pay large advances for new books.
It'll take a while before publishing collapses due to the availability of the same information via an LLM, but once this happens new books will get a lot more expensive for AI companies and a lot cheaper for everyone else.
I've been lead author and co-author on several books on programming. My first book sold tens of thousands of copies even though the writing was not just trash but a crime against humanity against anyone who had to read it. This was back in the day when the source code for the book was provided on request on a floppy, which was actually quite rigid because it was a Mac floppy.
My writing got better and better, and I got better publishers who actually hired editors, but the books sold fewer and fewer copies. Even crappy self-serving poorly written stackoverflow posts are often good enough. Then LLMs killed stackoverflow. Is that real ironic or Alanis ironic?
But I'm not holding my breath waiting for a book deal from OpenAI.
> My first book sold tens of thousands of copies even though the writing was not just trash but a crime against humanity against anyone who had to read it.
I’m sure I never read your book, but there is something cathartic about reading an admission like this. I have no doubt people got value out of your books, but I distinctly remember as a kid saving up to buy a couple expensive programming books from the bookstore and then being sorely disappointed in the way they were written. At the time I thought I was too young to understand adult writing, but when I went back to the books later as an adult it was obvious they were just written by amateurs. I still cherished them and learned a lot, but I will always remember the struggle of trying to follow along with what was probably some first-time writer’s attempt to learn how to write as they went along.
> But I'm not holding my breath waiting for a book deal from OpenAI.
Was this title eligible for class action status in the recent Anthropic scanning case? I know people who popped up on the list decades after writing obscure or forgotten technical titles. The estimated payout is $3k per title, usually split between publisher and author.
So... do the materials that it is printed on make it rare? Or is it the combination of ideas, paragraphs, and words that comprise the text make it rare and therefore valuable? If they're releasing digital versions of this text, i feel that is better off anyway.
They are only destroying the books because they are required to by copyright law. They obviously wouldn't do so if they were allowed to merely copy the book and preserve the original.
If it was $1 cheaper to destroy the books even if they didn't have to, they probably would. Storing books is expensive and selling it on means it could be picked up and scanned by a competitor.
Not true. To the extent that it is fair use to digitize a work for various purposes, it is also fair use to keep the original. It is only if they wanted to resell the original that they would have to delete their digitized copy. The reason they are destroying them because removing the binding is the most efficient way to scan them, they have no use for the originals after they have been scanned, and don't want to spend money storing them.
No, they aren't, for training AI, at least not based on anything but pure speculation.
The recent trial court decision that keeps being pointed to to support that:
(1) Found that for training AI, digitizing and copying works was fair use, period, with no requirement to destroy.
(2) For creating a centralized digital library for general use, digitizing works while destroying the hardcopy was fair use even with no intent to use them for training AI.
“And, the digitization of the books purchased in print form by Anthropic was
also a fair use but not for the same reason as applies to the training copies. Instead, it was a
fair use because all Anthropic did was replace the print copies it had purchased for its central
library with more convenient space-saving and searchable digital copies for its central
library — without adding new copies, creating new works, or redistributing existing copies.”
I’m not a lawyer, so it’s possible I’m missing something. But it seems to me like the ruling here implies fair use only holds so long as no net new copies, digital or otherwise, are created. If that’s the case, then it necessitates the destruction of the original.
It seems like most bookshop owners today are hoarders-turned-entrepreneur. They rent a storefront, adopt 2-5 cats, and they set up shop with their hoard of books so that they can hoard on a commercial level that would never be possible in a private home.
It also seems that American society has been swept away by the worship of books, rather than literary appreciation or promotion of education. "Banned Books Week" as the primary exhibit here. A tug-of-war over books that are supposedly "banned" when the verb itself has been twisted beyond recognition. I see posters and television shows and public service announcements that promote "Books!" and "Reading!" for no other purpose. Many books are trash and they will fill your head with garbage as sure as social media can, so why the indiscriminate worship of books?
It is conspicuous that, aside from established booksellers, and perhaps the LFL owners, many of the people crying out and moaning over book-destruction seem unwilling or unable to actually take those books in and store them. That is the point, that these books are unworthy of taking up storage space, which is an ongoing cost, and maintenance, and therefore, it is fiscally responsible to sell them, scan them, and destroy them. If you want to dedicate a room full of shelves to books that nobody will ever read, then go ahead and buy out your local bookseller! They will not stop you or cancel your order! Tell them you are saving those innocent books from the Big AI Bugaboo! They'll grovel at your feet!
Also now, I'm seeing small booksellers who feel "suspicious" or "skeptical" about large book orders. There was one Down Under who said "oh we got an order for over 70 and we can't physically handle that!" I feel like it's becoming a "shut up and take their money" situation. These booksellers, as professional business owners, should know that they may never liquidate stock, and this is their golden opportunity to simply get rid of some of that excess hoarded inventory.
But if you're merely cosplaying as a bookseller, and deep down you're a hoarder and a worshipper of books, you may really be reluctant to do transactions or earn money that sells away your books.
In more liberal sense, because you might be destroying something unique that later generations might actually like to see, the life rule that states "Don't be a dick" probably applies trumps even first sale doctrine.
If there was a public interest in these books then it was already attached, and it was an outrage for the sellers to exclusively possess these books, and it was an outrage to sell them to any other private party.
Does anyone else get pretty tired with 404s style of writing? I hate to make this comparison but it feels like a Fox News for the left.
I always struggle with their articles because it feels like they built it for rage bait on a topic and leave the other interesting topics out of it. It is an absolute shame for books to be destroyed but what does it really mean to be rare here? I know they kind of tried to differentiate but it sounds like this could be John Doe’s self help book that never sold well. If you ever are connected to a library you will start to realize how many books simply get thrown out or sold for nothing because nobody wants them.
To me the problem is partly the copyright law. I think it’s. Hard problem but I always lean more towards books having short copyright shelf lives and making it legal for digital copies to be shared after which I think would eliminate I good part of this problem. Not to mention 99% of the books published are probably garbage but that is highly subjective.
So it is bit of a meta rant but I think there a couple holes to go down that could be extremely interesting but they always write these informationally light articles. Like scrolling through a NYT visualization for just some shipping datapoints. Don’t really dig deep on anything and then end with a trust me bro these are rare books that Amazon is destroying for AI when I cannot be that upset with the amazons of the world. There are a lot of reasons for a business to digitize books, most books are worthless and it makes sense to cut the bindings for scanning. I would rather talk about how could you fix copyright to make this less an issue but is it even an issue with how many books get thrown out?
I agree with you. And it sucks, because 404 writes about interesting topics! But they're so negatively polarized it often obscures what's actually interesting.
> I always struggle with their articles because it feels like they built it for rage bait on a topic and leave the other interesting topics out of it.
I was going to say a very similar thing, and this is something I strongly disagree with when it comes to HN's moderation, and it goes like this:
The product is the outrage.
And HN should know better and mods should actively discourage, warn and prevent accounts (who karma farm, among other things) from even being able to post rage-bait articles. These aren't "hacker curiosities", they're just insipid bullshit. We wasted time and learned nothing.
Not really. It's not to my taste, but I've found their journalism to generally be lefty but focused on important events.
Like Amazon consuming and presumably destroying rare books should be enraging to everyone, regardless of political persuasion.
The difference is huge between 404 and Fox. Fox is out here trying to tell people there's a trans agenda, and that Biden was a lunatic leftist. They are just making up stories and publishing them because they know their audience engages. 404 definitely make editorial choices about which stories to pursue but I've largely found them to be grounded in real depictions of stuff that is happening.
Again, what rare books? Like they pointed to a book with low published volume or foreign language is “rare”. If you have ever been to annual library book sales you start to realize just how many books get thrown out. I can go buy a stack of books from the 1880s and 1890s on eBay right now for $25. $4 each. Most people don’t realize how worthless books are, even “rare” ones.
I even imagine that the market price for these “rare” books is helping filter out anything truly valuable and rare. It just reads as a rage bait tmz article. The quantity of used books including “rare” books that get thrown into the dump is astronomical.
Something can be "rare" while also having no value. One only needs to take a look at their local Facebook Marketplace listings to see this in action.
Rare invokes images of limited edition runs of well loved books, when in reality it's probably extremely outdated software guides, how-tos, technical manuals, etc.
> Like Amazon consuming and presumably destroying rare books should be enraging to everyone, regardless of political persuasion.
No, destroying collectible books would be a shame, not just any rare worthless books. But these are not collectible. The article tried to dance around it by saying maybe some books have a sentimental value to someone somewhere. But that doesn’t mean any library or collector wants it. Don’t fall for manufactured outrage!
Edit: Here’s an example of an extremely rare book. My great great grandfather published a book of sermons around 1920. That book has zero value to anyone other than my dad. Would I be outraged if it ended up at someone’s estate sale, then a used bookstore, and then an LLM consumed it to learn to read? No; I would have expected it to have been discarded by humans before the LLM even got to it. Most of what we leave behind is discarded.
follow the discussion around this on X, i think just a few minutes of research on this topic / reading past the headline you'll find out that this is pretty much a nothingburger
Why would that be a reasonable assumption? This is so blown out of proportion. The imagery invoked by the narrative is one of huge corporations destroying the final copies of literary treasures. That's just not happening.
Any bookseller worth their salt will ensure that a truly valuable book will not languish on their shelf for 1 minute longer than it takes to find a buyer willing to pay a fair price.
If AI-scanners are somehow bid-sniping bona fide collectors and wealthy aficionados, there may be cause for concern. But that is most certainly not happening here.
So is it moral of me to destroy a copy of "Windows 95 for Dummies"?
What if a person with small children and an elderly, incontinent pet with a penchant for peeing on books wants to buy it - can I sell this book to such a dangerous purchaser who might destroy it?
I love book as much as the next person but the hyperbole about "rare books" is absurd. Nobody is buying the Gutenburg Bible and destroying it for AI. The books in question are certainly not rare enough to be in museums - without titles there's no proof these are anything of real value.
To me, the hate against "destroying books" is kinda silly. The reason we're against destroying books is because the Nazis did it to destroy information.
This situation has absolutely no relation to that. This is the exact opposite, and the only reason this information isn't available publicly and is in risk of getting lost is copyright law.
TLDR: Amazon isn't the nazis in this story, copyright law is.
Making the verbatim content of the book lost forever is not equivalent to burning them? I'd even wager that the CO2 emissions of the scan,shred, and train process is even higher than burning.
But rich people want to violate copyright now! So judges ruled that it's now legal. But, because "the law is the same for everyone" they needed some excuse. NOT because, you know, obviously this demonstrates that very rich companies get to violate the law and you get hit by $30000 per infraction when you do it, when Anthropic ... doesn't even have to stop violating copyright when they do it.
Anyway ... this case is:
Bartz v. Anthropic PBC, No. 3:24-cv-05417-WHA, U.S. District Court for the Northern District of California, decided by Judge William Alsup
“the purchased print copy was destroyed and its digital replacement not redistributed, this was a fair use.”
Obviously this violates precedent, I had an internal LLM (probably a frontend for Claude or ChatGPT) find them (just like the convictions for file sharing in the 2000s required counter-to-the-law reasoning by judges, fair use was almost never accepted as a valid excuse, even when it obviously was, but of course Sony was a billion dollar company and needed to be in the right. In fact that this had to happen was explicitly given as a reason to create the DMCA)
Anyway, some precedents:
Hotaling v. Church of Jesus Christ of Latter-Day Saints, 118 F.3d 199 (4th Cir. 1997)
“Although the Church acknowledges that its sole remaining copy is not the one it originally acquired … it maintains that the remaining copy does not infringe Hotaling's copyright because it is a replacement copy…”
Atari, Inc. v. JS & A Group, Inc., 597 F. Supp. 5 (N.D. Ill. 1983)
... defendant sold a device for making backup copies of copyrighted Atari cartridges and argued that §117 permitted replacement/archival copying. The court rejected the broad replacement theory.
This very court has clearly declared that making a copy of a copyrighted work for replacement purposes is illegal, on multiple occasions.
I would like to point out that this isn't Anthropic's only extreme-WTF law violation. When the original judgement against them was made against them, they were forced to admit that using books to train models was illegal if acquired illegally AND THEN WERE ALLOWED TO KEEP DOING IT (thankfully the court never mentioned that part in the judgement so at least they can claim that was never decided when it becomes a huge problem in future cases as it obviously will). But that's not how this works. In my opinion Anthropic and OpenAI and everyone else need to at minimum take training material that was acquired in violation of copyright out of their training data unless and until they have a separate licensing agreement with the copyright holders. As long as Claude knows more about Harry Potter than is said in the promotional summaries it is obviously in violation.
Because of the copyright-filesharing court wars of the 2000s, which were also handled dishonestly by courts (whether we're talking US or EU courts), and the absurd copyright extensions, they had to now make some new excuse, and settled on this very sad, very destructive option. It's not even defensible legally, imho, but of course the biggest wallet must win. I don't understand. It's such a sad joke at this point, and it's not like the courts even still had credibility after the file sharing cases.
I wonder which sad excuse will be forthcoming from the courts when we have someone release a movie made by an AI model that is obviously a direct ripoff from some high-budget studio movie, and Disney needs to be protected from ... say ... "Scorched: Brothers of Aridelle" — In a vast desert kingdom, the royal brothers Elias and Anders grow up together, but Elias secretly possesses dangerous fire magic and isolates himself after accidentally hurting Anders as a child; years later, at Elias’s coronation, Anders announces his engagement to the seemingly charming Princess Hanna, provoking an argument that exposes Elias’s powers and sends him fleeing into the dunes, where he accidentally unleashes an endless heatwave that dries the kingdom’s wells and turns the capital into a furnace ...
I can think of a couple of solutions: 1. Glue a US $1 bill to the spine and take them to court when they destroy it. 2. Send them a license to use the book, with the condition that if they fail to return it within 30 days, they owe $1M. (Hey, if e-books can be licensed, why not physical books?)
In the end its not who can train the better model its who owns the better stack of training data. Acquiring old texts then destroying them makes their information proprietary.
How many AI companies are doing this? Say there are only 5 copies of a book left out there but 100 different companies want to shred it. Rare items could effectively vanish from the market. "Rare" meaning inclusive of high-quality items, disregard the bot opinion that mass book destruction is fine because most books are worthless anyway. And are any safeguards in place to make sure scanned books get saved in their original form for posterity before being recycled into trainingslop? Aside from Google's quasi-legal/ethical mass book scanning operation of course.
This is the citation needed that is missing from every report so far, including this one which deliberately refuses to reveal anything about these books.
You’d think if there were examples of actually valuable, rare books being shredded that the journalists would at least be able to name one such example. Instead it’s always vague posting about the destruction without ever naming any examples.
I think it’s because if they named some example titles, everyone would see that they don’t care about these books being shredded.
On the contrary these companies should be publicly listing the name/info of every book they destroy and use as training data.
If the books are really worthless as you say then their case would be proved transparently.
This article presumably had to maintain confidentiality to protect the seller who agreed to place a tracking device in the shipment.
I would say the burden of proof is on the companies destroying human cultural heritage en masse, not the handful of journalists calling for attention to the matter.
I do work adjacent to the AI book scan-shred pipeline. There are definitely significant books that aren't "Windows 95 for Dummies" which are getting down to single-digit remaining copies.
I just looked for one novel, which wasn't a fantastic book, but it is the first use of a pithy and fun phrase that is so ubiquitous that you'll probably read it a couple of times today. I argued with Claude, GPT and Gemini for ten minutes just now, even knowing the title of the book, to even prove the book exists. It took me years to find a copy originally and then I lost it in a move. I found one more copy today from a rare book seller, but it just sold (to Amazon?).
Is the book valuable? Not particularly, but I feel it's noteworthy and important. I don't want to name it either, because now I have some searches out and the next copy that pops up I'll scan and put on IA. There can only have been a few thousand copies originally published in 1947, it's only in hardcover. I know of a couple of other copies in private hands, so it's not zero copies, but it has to be single-digits.
I have one periodical issue that I know of only one other existing copy (Worthpoint only shows one copy ever sold in their database) and if you look on collector sites there is a blank because nobody even knows what the cover looks like. I can't explain it, since the publication routinely printed hundreds of thousands of copies of each issue, but here we are. Perhaps all the copies were withdrawn and pulped immediately after publication for some reason? It's in my scan pile, so I'll have it uploaded soon. Is it significant? Not hugely, but every other issue of this title has been scanned already, so it's scratching an itch to get this one done.
There's definitely rare stuff getting scanned and shredded. Someone in a comment above said it's not like Nazi book-burning since they were trying to destroy information. But it is like that if you consider there remain no other physical copies and all the electronic copies are locked up in a way that nobody can access except to trick an LLM to spit out a paraphrased copy from its training data.
TechCrunch is trying too hard to make a connection in the title there. Amazon was trying to make money then as it is now. Their selling of books at the beginning was no more principled than the destruction of books now.
This is the legal loophole that allows them to do what they need to do, and the benefit of doing it outweighs the modicum of outrage this title will generate.
Edit. Because I see my statement confused all the HN experts:
Anthropic's version of this was Project Panama [1]. The destruction of the original allows them to keep a digital copy in the AI training dataset. Quoting from the page:
> Judge William Alsup ruled that the destruction and digitization of legally purchased books constituted fair use
Is it a “loophole” that you can buy something and burn it? That’s how property ownership works. If the books were really so rare and valuable then the original owner should have given them to a museum rather than sell them as scrap. Yet no one who is complaining right now cared about them before Amazon got involved.
> Is it a “loophole” that you can buy something and burn it? That’s how property ownership works.
It sounds shocking if you're missing the context and only rely on the blogspam which is this TechCrunch article.
They aren't buying the books only to destroy them but to build an AI training dataset, which means they're making a digital copy. This becomes a copyright issue. The destruction is the legal loophole that allows them to keep the digital copy as the only copy in circulation and be considered fair use as decided by a judge (see below).
The concept of ownership means you can do anything you want with the object, the book in this case. Not with the content, like make or distribute copies. The scanning machines are cutting the spine of the book, feeding the pages to scanners, and then destroying the physical copy.
Anthropic's version of this was Project Panama [1]. Quoting from the page:
> Judge William Alsup ruled that the destruction and digitization of legally purchased books constituted fair use
No, that's not the loophole. The loophole is that a US federal judge recently ruled that scanning a book for the purpose of training an AI is "fair use" and hence legal _if and only if_ the book is destroyed in the process. If you retain the original, you would be creating an unauthorized copy, but if the original is destroyed, then it is permissable as format shift: https://fingfx.thomsonreuters.com/gfx/legaldocs/jnvwbgqlzpw/...
Sometimes original owner does not know that it is a first edition single copy. For them it is an old book. There have been many cases where owners did not know true value of their antiques.
all due respect but if they buy them they can do whatever they want with them. not like they are people and not like you have a right to them either. not even like you would have been able to acquire them yourself.
> As the bookseller who sold them told me, there are not many people in the world who would care about them in the same way people might care about the first edition of Oliver Twist, but that doesn’t mean they’re not valuable.
So books that would probably have ended up as trash. These AI training facilities are actually doing these book a service. Not only they are probably going to keep the scans safe (for future training), but having the book end up in an AI model may be the only way it is going to have any use at all.
Let's say for instance that the book in question is about woodworking, and it is not great, lots of mistakes and inaccuracies, unoriginal content, etc... except for a single thing, maybe a trick for making a certain measurement or something like that. Who would read such a crappy book for this single good trick he doesn't know is there, well, an computer will, computers process terabytes of crap without tiring and complaining, that's what they are for, and with a well designed LLM, that one trick may resurface, waiting for someone to ask about that specific measurement.
The problem here is not that rare books end up in AI training facilities, it is that these AI facilities are owned by for-profit companies keeping the data to themselves. These books should go to public libraries instead, for everyone to access, it would be better if these books weren't destroyed in to process too. But the question becomes: why didn't public libraries didn't do that in the fist place? And maybe in a more respectful way. The AI companies would just have had to license the database to libraries, probably simpler and cheaper than having their own scanning facilities.
To me, this mess is a failure of the copyright system. One one hand, large scale book digitization projects intended to preserve and make the original text accessible get lawsuits by publishers, while AI training is "fair use". It means we have built a system that encourages destroying rather than preserving these books!
Yes, we've inadvertently set up a system of incentives that were designed to preserve information and monetize it. And we've instead set up a set of incentives to make it scarce and destroy it.
I really have trouble getting worked up about this. "Rare books" is thrown around regularly but my gut feeling is that's not the case. These are used (often? always?) books and while I'm sure there is waste, in general they just want 1 of every book.
While I wish there was a repository of every book that was already digitized (it pains me this is the best solution), there isn't one and so I think this is not a real problem.
It'd be a different story if they had furnaces that ran only on rare books that they had to continually feed books to but that's not what's happening here. And that 1 destroyed copy will "live on" in a way that it otherwise might not.
It'll live on if they publish those scans or contribute them to a national archives or something. Proprietary data has a habit of being lost over time though.
They are literally not allowed to publish those scans.
Half the reason the books get trashed in this process is because the first sale doctrine keeps copyright from strangling all the freedom in this narrow area.
6 replies →
The problem is that we're ignorant of the true value of objects and we don't know what will be valuable in the future: https://en.wikipedia.org/wiki/Palimpsest
It seems a little short-sighted to destroy an artifact to get the text.
Perhaps the genetic material that remains in books from the people who handled them, or the pollen from plants in the environment that the book existed in will have value in the future, but we won't know what we lost because some people foolishly destroyed it in a bizarre quest to make AGI that the creators argue could potentially destroy humanity.
The more and more I read about these kinds of people the more I'm starting to realize that they're in the "here for a good time not a long time" group of people and those are the last people you want making long-term decisions.
5 replies →
We run a little library in front of our house. It's amazing how many people dump boxes of old books off in front of it hoping that they will find their way back into someone's collection.
The sad truth is the vast, vast majority of printed literature is neither interesting nor useful. People are not dumping off stacks of Umberto Eco. We frankly have to toss a lot of awful cookbooks, self-help books, trashy mass-market "novels", and sketchy religious works. As it is, even the stuff that makes it to the library is not very impressive.
Imagine if we had bad cookbooks, cheap popular novels and stories, tracts on diet and self help from old Rome, or old eras in China, or the equivalent from ages before that.
It's all interesting for something, even if it's just a meta analysis of culture during a certain period or what kind of trashy romance novels were popular in 198X. At least in my view.
4 replies →
Do we know if they actually buy one copy of each book here? Or are they indiscriminately buying books in bulk and just processing all of them?
The latter seems inefficient, so my first assumption would be that they would avoid that. But while I suspect they'd check if they know the book before scanning, I could imagine them not caring that much before buying them and just focus on volume.
> they just want 1 of every book.
That's the point I come to.
Books are not original manuscripts. Even in low volume cases, they are usually printed hundreds of times. (And usually low volume works aren't all that great...hence the low demand.)
That is a fraction of a percent for books that at some level weren't all that wanted.
Agree. Couldn't care less. There are some neat aspects to "rare books" but overall quite insignificant.
Rare books tend to be out of copyright, true.
The Embassy of the Free Mind (https://www.embassyofthefreemind.com) is a rare book library in Amsterdam that is scanning books the old fashioned way… leading to https://SourceLibrary.org — a collection of over 5,000 books from the renaissance that have never been translated before. Consider donating, if this is a topic you care about!
The length of this article and the unnecessary visualizations felt like a waste after getting to the end and discovering they won’t reveal anything about the books that were scanned.
The closest they got was admitting that the rare books weren’t anything that someone might care about in the sense that people assume when we hear “rare books”
> As the bookseller who sold them told me, there are not many people in the world who would care about them in the same way people might care about the first edition of Oliver Twist, but that doesn’t mean they’re not valuable.
Okay? But then why exactly where they considered valuable enough to warrant an entire article about them going to a book scanning facility? Without revealing anything about these books I have no idea if they were classic literary works that were underappreciated, or if this was some old guide about How to Use Microsoft Office 97.
There was a more balanced take on Twitter (which I’m unable to find again, because Twitter) from a book seller who said it was more of the latter type: Books that were rare because they were no longer in demand and most everyone had thrown their copies away. Some parts of the media are doing backflips to try to imply that these are cherished literary classics being fed into the shredder to deprive humanity of something valuable, but the book seller seemed happy to be making sales for useless old books that no human was interested in buying.
As if the AI industry is only shredding "crappy" books.
I recently bought a rare book. It was a boat design book written by a famous yacht designer in the 1940s, but it's been out of print for decades and I had to pay $150 for a "fair" copy with missing dust jacket.
I have an interest in older technology and the old ways of doing things (for instance, how do you lubricate the mast of a gaff rigged boat so the gaff jaws don't jam?) I've often found myself reading very old books that have been out of print for a century.
Sometimes I read those books at libraries and I've been the only person to check them out in years (I started doing this back when they still stamped the return date on a card so you could see when it was checked out). Now most of those books have been disposed of by libraries due to yield management software and I've ended up with some of them in my personal collection, but people like me can only save a tiny sliver when most of them are being bought in mass by Sam Altman. When they're gone the knowledge in them is also gone.
We are burning the library of Alexandria and the HN consensus is "those books probably weren't saving anyway."
> As if the AI industry is only shredding "crappy" books.
Do you, or anyone else, have any source suggesting that they’re buying highly valuable rare books and shredding them?
The kind of $150 rare book that you had to buy from a specialty collector who graded it is in a completely different category. You’re thinking of “rare books” in the historically rare, valuable, and collectible category.
The book sellers shipping off orders of 1000s of books at a time to these facilities are calling the books “rare” because they may only have 1 copy, not because it’s a collectible with a high price tag.
9 replies →
Please scan this book and put it into Anna's library, or keep it for the future.
I would be also interested in participating in your costs.
Im with you, my take-away from the article was of sadness for all the books that are going extinct because of this.
Real humans sharing their unique knowledge, packaged in a book.
1 reply →
It would be nice if your local library hadn't sold their copy but I think your beef is with the libraries, or perhaps with the politicians who failed to fund them in that case.
Actually I expect that a copy of most books probably is still available in a copyright library but it might not be easy to access.
https://runtimewire.com/article/anthropic-settles-book-pirac...
https://www.gadgetreview.com/we-dont-want-it-to-be-known-ins...
https://www.irishtimes.com/world/europe/2026/08/10/a-mysteri...
https://www.bbc.co.uk/news/articles/cp3rprx2wl4o
https://dallasexpress.com/national/the-vanishing-page-ai-fir...
https://fortune.com/2026/07/31/dutch-bookseller-ai-spam-phis...
Weeding (the library term of art for selecting works for deacquisition) is a very import part of collection management. It's actually regionally and nationally coordinated so the interlibary loan network does not throw away the last copies.
https://cdlib.org/west/ https://papr.crl.edu https://eastlibraries.org
2 replies →
Once you start to realize just many books got thrown out annually before “AI” you have less sympathy. Libraries are one of the biggest contributors to this because there is simply too many books that nobody cares about.
I suspect folks are over weighting how much knowledge is being destroyed in these books. If someone actually quantified it that would be amazing but as someone who started going to used book sales at a very young age I just have no sympathy. Most books are worthless. I don’t mean that from a text perspective either.
10 replies →
Can we copy the books? It seems like in an effort to monetize writing, we've created a system that incentivizes destroying it.
> or if this was some old guide about How to Use Microsoft Office 97.
Materials like that may be very valuable to software / tech / HCI archeologists soon.
I can't get the tools or local know-how to straighten my scythe blade in a country where every cottage had a scythe less than 100 years ago, with the last scythe-native generation rapidly dying out.
Fast forward a civilizational collapse and that M$ Office 97 for Dummies might be as groundbreaking as a Guide to Using Roman Concrete
This is why we read the comments first :)
Thank you.
Anyone who is into books eventually finds out that we're permanently losing them all the time. Like they're thrown out, and lost forever. University libraries throwing out huge collections of out-of-print material to make room for new books, study spaces, even cafés. Municipal libraries turning over their collections. Books that never made it to libraries going out of print, tossed in the trash after yard sales.
The AI companies digesting this stuff is a net win for humanity. And I'm not a fanboy! Ideally they'd upload them to Anna's archive too, but even if they keep it private forever, at least these books live on in some way in the model weights. Thats better than a landfill.
It would be really great if the companies who are doing this would commit to placing the scanned files into a public trust that would coordinate with organizations like the Gutenberg project to ensure that the scanned materials enter the public domain on schedule. Publishing encrypted archives with the keys in escrow would be a good first step.
IMO that would go a long way to resolve any concerns about losing books. I still don't like the idea of extremely hard to find or last prints being actually destroyed for this, but it certainly makes it more palatable.
15 replies →
I would feel much better about this process if they were uploaded and if it were framed as a knowledge preservation project. This would only slightly increase the cost of the project, but have a huge impact on its perception and its net positive impact.
Of course, actually benefiting humanity is only a minor, indirect concern for investors.
3 replies →
If or when they go bankrupt or reach agi, they will just delete them. I hope Anna's archive already has them anyways. Apparently openai's newest unreleased model, gpt 6, is capable of continuous training at Inference time, like a person is. That might be enough to delete them
6 replies →
That is the problem: they are digesting it. They are not creating a new kind of library, where you could say, show me the text of "How to Fix Your Ice Cream Problems". (An actual book I own) It is not being done for our future reference.
3 replies →
> Ideally they'd upload them to Anna's archive too
Not only can they not do that, they must scan physical copies because they are forbidden from using digital pirated copies from sources like this.
Anthropic had a big settlement because they were caught using downloaded digital copies. As a response they’ve ramped up their book scanning and others have followed.
1 reply →
This is the right answer but the reason they can’t upload the scans is copyright law as demonstrated by Google having to settle with the publishers and allow them to remove their books and limit free access to 20% of text. The AI companies are essentially compressing the information in a huge swath of books that would otherwise be headed to landfill and making them 1000x more accessible. This is unquestionably one of those instances where capitalism is taking money from rich investors and benefiting the 99%.
This feels like a manufactured controversy. What difference does it make to me what someone does with a book after they buy it? It's effectively unavailable to me regardless of what they do. If people are really concerned about these "rare" books, they should lobby the copyright owners to release them online or print more copies.
> What difference does it make to me what someone does with a book after they buy it?
I think this is a bit of a myopic take. It's like saying it's my property I can do what I want, yet there exists designated historic homes or neighborhoods that are deemed to have cultural value where modifications do in fact need to be approved.
I'm with you on the "rare" part. If people are thinking about 70+ year old documents or ancient manuscripts, I doubt that's what AI is being trained on and is being destroyed, but it's reasonable that people find _that_ idea distasteful.
You can say it's manufactured but if these companies ignore this criticism, it's just another way AI companies are committed to losing the public.
> It's like saying it's my property I can do what I want, yet there exists designated historic homes or neighborhoods that are deemed to have cultural value where modifications do in fact need to be approved.
Yes, and when you buy them with that designation, you know what you're getting into. The problem occurs when you already own it, and some group is trying to get it labeled as historic, which will add to your burden and limit what you can do with it.
When I was in a small town, this was actually weaponized. A hotel owner was trying to get another hotel categorized as "historic" and had rallied a lot of people behind his cause. He had a case - the hotel did have some claim to being the "first" in some category or other. But really, he was doing it because it was a competitor. The "historic" hotel owner had to spend a lot of money to fight the cause, because being labeled historic would prevent him from performing various upgrades, making the hotel less attractive to customers (he was already not getting many customers).
What difference does it make to me if someone shoots the last bison? I wasn't getting to eat it either way.
If people are really concerned about bison, they should lobby gamekeepers to release photos of them.
https://en.wikipedia.org/wiki/American_bison#/media/File:Bis...
Replace bison with some random animal no one's ever heard about, and photos with literal clones of the animal, and you've got a much better metaphor. Not to forget that in this world, brand new animals are created every day.
2 replies →
A digital copy of a book is identical in value to a printed copy.
3 replies →
If there’s only 3 copies in existence, and everyone on the frontier wants it in their corpus, what do you think will happen?
This is a problem whether it's AI companies buying them or three random dudes. The real solution is to get the copyright owners to keep distributing copies, or to change copyright law.
3 replies →
If there are only 3 copies, then each copy is going to cost thousands of dollars.
Sample:
> Thou shalt commit adultery.
https://en.wikipedia.org/wiki/Wicked_Bible
I think the term "rare" might be a bit loaded. DOES it mean "only a handful in existence" or more like "1000 copies?" And at any rate, at least by scanning the book they're theoretically making its contents available to the public. What are the other people who own these books doing besides having them sit on a shelf?
4 replies →
> what do you think will happen?
They raise the price and print more copies?
11 replies →
It's better value to me that someone is digitizing and training with books then they just languish somewhere in perpetuity.
> What difference does it make to me
It's the scale that matters as first, and secondly, most people don't shred their books after reading them once or twice. This is just beyond words.
> We’re not revealing the titles of the books included in the shipment we tracked, but they are rare, meaning there are not many copies of them in circulation. Sometimes that’s because not many copies of them were ever printed, and sometimes because they are in a foreign language not many people speak.
Not even the title of one of those rare books?
> The seller agreed to put an Apple AirTag provided by 404 Media in one of the books included in this order so we could see where the book was going. And by extension, which company, AI or otherwise, was behind this massive order.
It would seem the redaction of the rare titles is a way to avoid de-anonymization and subsequent harm to the business of the seller who agreed to place a tracker in one of the books. That being said, maybe they could have chosen a better methodology which would have allowed the disclosure of the title, although ultimately I’m not sure the title matters too much outside of their claim they were “rare”.
The issue is if rare is simply "not many in circulation", that's very very different from rare being "it is a unique artifact". The former kind of rarity doesn't really present an existential end, whereas the latter does.
7 replies →
Makes you wonder how many packets with rare books.containing airtags said recipient receives. My guess is: one so far.
1 reply →
And even then, the "rare" qualifier is not needed here. What Amazon/LLM companies are doing is amoral. This is intellectual piracy (not in the "copyright infringement" sense) at the highest level. Stealing and centralizing the accumulation of human knowledge to eventually rob us all and put all power into the hands in the hands of a handful of people, who are not benevolent.
10 replies →
It doesn't really matter. The article suggests that they are selecting and tracking books by ISBN, which means: books that have been published or at least reprinted in the last 50 years or so. And they are trying to get as much of that set as they can, regardless of quantity, quality, or any other consideration. Which means that older books printed before ISBNs became common may be relatively safe, at least from Amazon. And some books just don't come up on the used market very often.
(small historical irony: when Amazon first started selling books, they used the Books in Print database, which included a lot of books not actually in print.)
As I understand it, "rare" in this context could mean anything—even a washing machine manual from the 1980s...
Rare as in your grandfather's John Deere manual from 1982, not rare as in a test print run of The Great Gatsby. The number of books ingested by these AI companies is a drop in the bucket compared to old books destroyed every year through normal means.
> The number of books ingested by these AI companies is a drop in the bucket compared to old books destroyed every year through normal means.
Citations? Also what exactly are these 'normal means'? As a bibliophile who loves scouring used book stores for out-of-print titles this is a topic I'm very interested in.
Here's EPA data from as recent as 2018: https://www.epa.gov/facts-and-figures-about-materials-waste-...
EPA estimates that hundreds of thousands of tons of books are landfilled or recycled every year. That's just the United States. It seems likely that the worldwide figure is in the millions of tons.
> Also what exactly are these 'normal means'?
Normal means is throwing it in the recycle bin. Especially for stuff like a 1982 John Deere manual. I've never donated an old appliance's manual to the library. Have you?
Amongst my friends, I'm one of the rare folks who donates books to the library. Most people just trash them. And I know the library only wants them to try to sell them in their book sales (or online) so they can get money. Almost nothing one donates to a library actually ends up on the library shelves.
"rare" is used in these headlines/articles to incite and generate clicks
The scariest thing for me about companies not caring even minimally about conservation is that when AI gets more powerful than humanity (which is clearly a when not an if, even if there's a lot of uncertainty and differences in opinion about the time horizon here), I want to hope that AI will care more about conservation of human people.
So far from how I see how powerful organizations work, I'm not as certain as I would like to be.
> AI gets more powerful than humanity (which is clearly a when not an if, even if there's a lot of uncertainty and differences in opinion about the time horizon here)
People say this all the time, but so far nothing has convinced me it's true.
LLM development has more or less plateaued, and the current boundaries are very real - energy, resources, capital.
At this point we're talking about marginal improvements against the same asymptotes of all technological innovations.
So, if a corporation can suck up entire books to teach their machines how to think using the information from those books, can we (all humans) join a single corporation that provides all books to its employees? Just need one copy of each and we'll make that copy digitally available for our employees so they can learn from and utilize the knowledge from the books. It's not copyright infringement, they're employees.
Almost like… a library?
This is what the copyright laws dictate no ?
Bias disclaimer: Amazon is my current employer, but I don't work on AI or anything else mentioned in the article.
Yes, this is a result of copyright laws. The other commenters are wrong/uninformed.
If it was up to the companies training LLMs, they wouldn't destroy the books: It's a waste of company resources, it's needlessly destructive/evil, it generates bad PR, etc etc. There are essentially zero advantages, other than it is what is required under US copyright law (or at least, it is what their highly paid lawyers believe is required under US copyright law).
Isn’t it being destroyed because it makes the scanning process easier?
8 replies →
They could just not do the evil thing.
(This is why I will never be a billionaire)
Not to my understanding. To begin with, it's far from given that "rare" books are all covered by copyright. But if they are, it's at best murky: whether you destroy the original doesn't really have anything to do with what you're doing with scanned contents. The scanned contents themselves may be inherently a copyright issue, regardless of destroying the original. The actual trained model has separate arguments more in its favor, so if no scanned contents exist - IE the data is read once for training and not stored or saved, they have a better argument. But in that case the destruction is totally disconnected from copyright, as they'd be totally okay to rescan the material.
Those laws are lobbied for by large corporations, these are not just laws that exist outside of that context. They can also be changed, or Amazon could just incur the fines.
Large corporations will move fast and break things when it’s convenient; they don’t care much about the law - just about profit.
i would assume that would come into play if they were uploading scans of the books? there must be some gray area where training like this doesnt apply to that. or ya know, just do it and face the consequences later because you already have the data and know that ai obsessed government will just shrug their shoulders.
No. A copy is still a copy even if you destroy the original.
Until about a year ago this would have been a reasonable and respectable argument, but at least in California you are arguing against current legal precedent: https://fingfx.thomsonreuters.com/gfx/legaldocs/jnvwbgqlzpw/...
2 replies →
Discussions:
2 days ago https://news.ycombinator.com/item?id=49068738
I support physical media and doing whatever you want with said physical media. If you want to overpay for a copy of Sharepoint 2007 For Dummies and destroy it, knock yourself out. Just because media has been printed and is "rare" doesn't mean it has any practical value.
You could easily write an article about how Goodwill and the Salvation Army dump millions of "rare" books that no one would ever buy for $0.50. And thrift stores will dump multiple copies of said books. AI companies would only ever need one!
The only thing I disagree with is AI companies feeling entitled to freely use every piece of copywritten work without restriction, but that applies just as much to web scraping, pirated books, etc.
404media is paywalled for this article now.
Here’s Tom’s version:
https://www.tomshardware.com/tech-industry/artificial-intell...
Why is the textual data bound up in these paper books worth the trouble?
These LLM training runs have already ingested essentially the whole public internet. What marginal value is to be gained from scanning and destroying obscure books?
For me the biggest functional issues with LLMs don't seem to have any connection with "I wish they had read this obscure community cookbook from 1946". Is that going to get Claude to stop saying "honestly" to me? Is it going to get LLMs to stop making up sources that don't exist? What is in it for Amazon or any LLM company to chase more obscure data like this.
Frontier researchers have found that dumping more and more data into training is effective at improving LLM capabilities. Nobody has a gears-level understanding of how training on some particular kind of data leads to some particular behaviors, so they generally take the attitude that more is better.
I have to agree; even if they are getting regular 20th century out-of-print books, is that going to add a significant percentage to their training data?
I can only think of it being a 'low-background steel' situation where they want to locate original, non-digitized text for validation or knowledge bases.
I'm not an expert in LLM training, but I think we can all agree that the writing on the Internet is generally very low-quality and surface level compared to the depth of books. Most books in the past were even edited by a separate person from the writer!
Humans manage to be pretty intelligent with only reading perhaps a few thousand books in their lifetime. It seems unlikely that AGI will appear but only once it has read that last out of print 1983 book on knitting patterns.
Granted that that is indeed the case, surely the quantity of high-quality writing available online (including things like Project Gutenberg) still dwarfs the quantity remaining in obscure unscanned books.
We really need to be thinking about the long-term implications of what Amazon, Anthropic, OpenAI, and others are doing here.
Because it's not just one company doing this, it's dozens if not hundreds of them, all with huge budgets and all competing over the same dwindling supply of older books. What those talking about library disposals and similar measures are missing is the sheer scale that destructive book scanning for AI ingestion is operating with. It's unprecedented.
It's not impossible that in a few years the number of older physical books for sale anywhere will plummet to almost nothing and that in many cases that'll include the last copies anywhere of particular titles. Second-hand book stores will close, and a ton of niche knowledge might be lost forever. It'll also all have been done in relative secret, without any transparency or public record kept of what was lost.
Worse, it's something that can only be done once. If our societies don't do something about this now there isn't going to be a chance for a do-over. Once a physical book is destroyed it's gone forever, and we'll be lucky if there's a digital copy left. But even then digital copies don't provide the same level of forensic verifiability that physical copies do. Future researchers looking for material that might've been contained in books like these will be out of luck.
If our societies do nothing to stop this, we'll all be poorer for it.
I don't understand why Amazon still bothers to blow money on their in-house AI effort (outside AWS Bedrock). What have all the scraping, scanning and training led to? Nobody is using their models.
Complete stab in the dark: AWS "training as a service" (further split into multiple microservices) for companies looking to train their own models.
AWS has nova and titan models that nobody seems to use
1 reply →
The dataset is valuable on its own without the model.
A lot of the rare books seem to be books that are most relevant to specific professions within specific countries. I would expect university or national libraries in those countries to have copies of those books, but they often do not digitize them because controlled digital lending is seen as a legal grey area at best. So these ai companies are doing digitization for them and not publishing it, duplicating work for these libraries once they do decide to digitize.
Many other rare books are just rephrasings of other books on a given topic and this can be done with synthetic data generation now. Same goes for websites
Honestly this feels like a US exclusive problem. Like shipping from overseas would be way too expensive for the cheap books they want. The Chinese just use Anna's archive instead
A couple thoughts:
Would that an enhancement to the Library of Congress or such?
A couple answers…
Why do you think for even a moment that they would buy only one copy, rather than every available copy of something rare and difficult to get? Anyone doing this stands to gain from destroying the only copies of something that has scarcity, to stop competitors from ever getting access to it. That's a serious concern.
Why would they do that? What is this 'should' to a company that would do this in the first place? Are you not suggesting the opposite of what they aim to produce?
That seems really speculative. I think you're assuming the maximum possible evil intent for no reason, mostly because you hate AI and anyone involved in it.
While we can quibble (and are) over the details of this particular occurrence, I'm beginning to take the opinion that used/rare booksellers need to implement some form of "KYC" to help guard against the epistemic threats posed by this practice.
The booksellers are probably so ecstatic to get some value out of books otherwise destined for the landfill that they wouldn't lift a finger for "epistemic threats".
I remember when google was scanning a bunch of rare books, I mean they might still be doing that? Either way, that was cool.
I have a few "rare books" and have read many, you'd be suprised at what is publicly available on google books since like ~2010ish.
Years ago I saw an article on this. For Google Books, they had two processes.
The first was destructive. This was for mainstream books currently being published so they had no value. It's (I believe) where you cut off the spine and scan the pages.
For rarer books, there was a non-destructive process. Basically the book was opened to each page and scanned. This was slower but didn't destroy the book.
I don't understand why these companies haven't just licensed the scans Google has already done. Why is each company doing this rather than just scanning the books once and sharing the scans?
> I don't understand why these companies haven't just licensed the scans Google has already done.
Because for the vast majority of the books, Google doesn't have the legal right to license those scans. It would be legal if the books were out of copyright, but despite the connotation of "rare books" those generally aren't the books we're talking about. Further, in the cases where Google didn't destroy the original of the book they scanned, their scanned copy may be considered infringing under the new standard, so Google doesn't want call undue attention to what they have.
1 reply →
Here’s a clip of The Internet Archive’s nondestructive process [39 seconds]:
https://nitter.net/internetarchive/status/135809098218971955...
Wonder the additional cost to add automated page turning. More than Bezos can afford, impoverished chap he is.
If I remember correctly, it wasn't a distinction between common and rare books.
It was a distinction between library books and others. They partnered with libraries, and obviously, libraries didn't want their books destroyed, so they devised a non-destructive book scanner. Some of the books they scanned were indeed rare, but rare or not, you don't destroy books you borrowed from a library!
Google probably doesn't have the scan, and now, if they do, it might be proprietary.
According to the article, the scanning facility has been there for a while, so it may well be that they have been destructively scanning books for a while, even before AI. It makes sense for the major book retailer to have digital copies of every book so that it can begin selling those as soon as it can negotiate the rights (if there are any).
In any case, all the indignation about destructive digitization misses the point that rare books takings space in a warehouse for years without being bought will eventually be destroyed anyway.
Something I've been thinking about is that a physical book is a physical artifact, and it's age and authenticity can be verified by treating it as such.
In contrast a digital book has no such marks. It's impossible to tell if it was redacted to hide inconvenient passages, or even completely rewritten or fabricated (which can now easily be done at scale with AI). If the physical book is a primary source then the digital version must be considered a secondary source: A potentially biased retelling.
For this reason, a digital book can never be a perfect substitute for the physical book it was created from.
I watched Short Circuit yesterday. All I can think now is "need input".
Am I too old now, expecting someone to make a Rainbows End reference? Vernor Vinge predicted this 20 years ago.
(Also the person who coined Singularity, though Ray Kurzweil really wanted everyone to think it was his idea.)
Yep, unpleasantly close prediction. Probably shouldn't give them any ideas.
That is certainly an interesting case for thinking things but not typing or saying them out loud.
Problem is, they all read the same books we do.
Great article and work to reverse engineer who was mass buying used books.
>It'll live on if they publish those scans or contribute them to a national archives or something
I'm not really aware of many benevolent acts Amazon has taken in the past decade. Are you?
I'm less concerned with how these rare books didn't rot in large lots of unused books and more concerned with whether or not this helps preserve the content longer.
I seriously doubt any of the companies doing this will preserve the scans for very long. It costs money to store stuff, and none of these companies are have any concern for anyone who isn't them, so they're not going to spend the money or lift a finger unless they work out a way to make it profitable.
archive.org estimates it costs $2/GB to store data in perpetuity (well, at least until decades of storage scaling trends stop of course). That's probably less than it costs to acquire and scan the data in, I'd be surprised if they just tossed it out at the end when it's so cheap to store. The same thing happened at the healthcare data lakes I worked on where the trend switched from the usual "how long are we legally required to store this information" to "how long can we legally hold this information".
The content is preserved obviously, and I hope that sometime in the future it will be made available in its original form.
What worries me about the trends is inevitable sanitization of content or straight out falsification.
>What worries me about the trends is inevitable sanitization of content or straight out falsification.
That sounds really speculative, and not inevitable at all.
I think you're just trying to invent things to be worried about because you don't like AI and don't trust AI companies.
1 reply →
... because now it can be done at scale.
The content is never made available. It would be illegal to.
It will be legal to share in some decades. Not that Amazon will bother, though.
I don't expect an amazon.com/checkoutrarebooks page but I also don't really think the legality of sharing the content has been much of a concern for these companies either. Ironically, one of the few things Meta got in trouble about with building their AI was helping share the training data.
https://web.archive.org/web/20260817141552if_/https://www.40...
I look forward to the day that AI companies pay large advances for new books.
It'll take a while before publishing collapses due to the availability of the same information via an LLM, but once this happens new books will get a lot more expensive for AI companies and a lot cheaper for everyone else.
I've been lead author and co-author on several books on programming. My first book sold tens of thousands of copies even though the writing was not just trash but a crime against humanity against anyone who had to read it. This was back in the day when the source code for the book was provided on request on a floppy, which was actually quite rigid because it was a Mac floppy.
My writing got better and better, and I got better publishers who actually hired editors, but the books sold fewer and fewer copies. Even crappy self-serving poorly written stackoverflow posts are often good enough. Then LLMs killed stackoverflow. Is that real ironic or Alanis ironic?
But I'm not holding my breath waiting for a book deal from OpenAI.
> My first book sold tens of thousands of copies even though the writing was not just trash but a crime against humanity against anyone who had to read it.
I’m sure I never read your book, but there is something cathartic about reading an admission like this. I have no doubt people got value out of your books, but I distinctly remember as a kid saving up to buy a couple expensive programming books from the bookstore and then being sorely disappointed in the way they were written. At the time I thought I was too young to understand adult writing, but when I went back to the books later as an adult it was obvious they were just written by amateurs. I still cherished them and learned a lot, but I will always remember the struggle of trying to follow along with what was probably some first-time writer’s attempt to learn how to write as they went along.
> Then LLMs killed stackoverflow.
I think the common consensus is that stackoverflow killed stackoverflow, quite a few years before LLMs became entrenched.
> But I'm not holding my breath waiting for a book deal from OpenAI.
Was this title eligible for class action status in the recent Anthropic scanning case? I know people who popped up on the list decades after writing obscure or forgotten technical titles. The estimated payout is $3k per title, usually split between publisher and author.
https://www.authorsalliance.org/2025/09/07/the-anthropic-set...
2 replies →
So... do the materials that it is printed on make it rare? Or is it the combination of ideas, paragraphs, and words that comprise the text make it rare and therefore valuable? If they're releasing digital versions of this text, i feel that is better off anyway.
Well, I hate to be the one to break it to you but they're not releasing digital versions, they're just destroying them.
They are only destroying the books because they are required to by copyright law. They obviously wouldn't do so if they were allowed to merely copy the book and preserve the original.
If it was $1 cheaper to destroy the books even if they didn't have to, they probably would. Storing books is expensive and selling it on means it could be picked up and scanned by a competitor.
Not true. To the extent that it is fair use to digitize a work for various purposes, it is also fair use to keep the original. It is only if they wanted to resell the original that they would have to delete their digitized copy. The reason they are destroying them because removing the binding is the most efficient way to scan them, they have no use for the originals after they have been scanned, and don't want to spend money storing them.
No, they aren't, for training AI, at least not based on anything but pure speculation.
The recent trial court decision that keeps being pointed to to support that:
(1) Found that for training AI, digitizing and copying works was fair use, period, with no requirement to destroy.
(2) For creating a centralized digital library for general use, digitizing works while destroying the hardcopy was fair use even with no intent to use them for training AI.
There's no copyright law that requires the owner of a copy of a book to destroy it. What are you talking about?
From the ruling of Bartz v. Anthropic:
“And, the digitization of the books purchased in print form by Anthropic was also a fair use but not for the same reason as applies to the training copies. Instead, it was a fair use because all Anthropic did was replace the print copies it had purchased for its central library with more convenient space-saving and searchable digital copies for its central library — without adding new copies, creating new works, or redistributing existing copies.”
I’m not a lawyer, so it’s possible I’m missing something. But it seems to me like the ruling here implies fair use only holds so long as no net new copies, digital or otherwise, are created. If that’s the case, then it necessitates the destruction of the original.
1 reply →
It seems like most bookshop owners today are hoarders-turned-entrepreneur. They rent a storefront, adopt 2-5 cats, and they set up shop with their hoard of books so that they can hoard on a commercial level that would never be possible in a private home.
It also seems that American society has been swept away by the worship of books, rather than literary appreciation or promotion of education. "Banned Books Week" as the primary exhibit here. A tug-of-war over books that are supposedly "banned" when the verb itself has been twisted beyond recognition. I see posters and television shows and public service announcements that promote "Books!" and "Reading!" for no other purpose. Many books are trash and they will fill your head with garbage as sure as social media can, so why the indiscriminate worship of books?
It is conspicuous that, aside from established booksellers, and perhaps the LFL owners, many of the people crying out and moaning over book-destruction seem unwilling or unable to actually take those books in and store them. That is the point, that these books are unworthy of taking up storage space, which is an ongoing cost, and maintenance, and therefore, it is fiscally responsible to sell them, scan them, and destroy them. If you want to dedicate a room full of shelves to books that nobody will ever read, then go ahead and buy out your local bookseller! They will not stop you or cancel your order! Tell them you are saving those innocent books from the Big AI Bugaboo! They'll grovel at your feet!
Also now, I'm seeing small booksellers who feel "suspicious" or "skeptical" about large book orders. There was one Down Under who said "oh we got an order for over 70 and we can't physically handle that!" I feel like it's becoming a "shut up and take their money" situation. These booksellers, as professional business owners, should know that they may never liquidate stock, and this is their golden opportunity to simply get rid of some of that excess hoarded inventory.
But if you're merely cosplaying as a bookseller, and deep down you're a hoarder and a worshipper of books, you may really be reluctant to do transactions or earn money that sells away your books.
I remember when hackers believed in the doctrine of first sale. You can do whatever with the stuff you own.
In a strictly libertarian sense, sure.
In more liberal sense, because you might be destroying something unique that later generations might actually like to see, the life rule that states "Don't be a dick" probably applies trumps even first sale doctrine.
If there was a public interest in these books then it was already attached, and it was an outrage for the sellers to exclusively possess these books, and it was an outrage to sell them to any other private party.
More and more, I am glad that I stopped giving Amazon my (formerly) enormous amount of business.
The headline implies at some level that Amazon loves or cares about ..... stuff.
Nothing but a drawing of an alligator eating a book. Where is the article?
Link to URL? There is no article at the above link.
Does anyone else get pretty tired with 404s style of writing? I hate to make this comparison but it feels like a Fox News for the left.
I always struggle with their articles because it feels like they built it for rage bait on a topic and leave the other interesting topics out of it. It is an absolute shame for books to be destroyed but what does it really mean to be rare here? I know they kind of tried to differentiate but it sounds like this could be John Doe’s self help book that never sold well. If you ever are connected to a library you will start to realize how many books simply get thrown out or sold for nothing because nobody wants them.
To me the problem is partly the copyright law. I think it’s. Hard problem but I always lean more towards books having short copyright shelf lives and making it legal for digital copies to be shared after which I think would eliminate I good part of this problem. Not to mention 99% of the books published are probably garbage but that is highly subjective.
So it is bit of a meta rant but I think there a couple holes to go down that could be extremely interesting but they always write these informationally light articles. Like scrolling through a NYT visualization for just some shipping datapoints. Don’t really dig deep on anything and then end with a trust me bro these are rare books that Amazon is destroying for AI when I cannot be that upset with the amazons of the world. There are a lot of reasons for a business to digitize books, most books are worthless and it makes sense to cut the bindings for scanning. I would rather talk about how could you fix copyright to make this less an issue but is it even an issue with how many books get thrown out?
I agree with you. And it sucks, because 404 writes about interesting topics! But they're so negatively polarized it often obscures what's actually interesting.
> I always struggle with their articles because it feels like they built it for rage bait on a topic and leave the other interesting topics out of it.
I was going to say a very similar thing, and this is something I strongly disagree with when it comes to HN's moderation, and it goes like this:
The product is the outrage.
And HN should know better and mods should actively discourage, warn and prevent accounts (who karma farm, among other things) from even being able to post rage-bait articles. These aren't "hacker curiosities", they're just insipid bullshit. We wasted time and learned nothing.
Not really. It's not to my taste, but I've found their journalism to generally be lefty but focused on important events.
Like Amazon consuming and presumably destroying rare books should be enraging to everyone, regardless of political persuasion.
The difference is huge between 404 and Fox. Fox is out here trying to tell people there's a trans agenda, and that Biden was a lunatic leftist. They are just making up stories and publishing them because they know their audience engages. 404 definitely make editorial choices about which stories to pursue but I've largely found them to be grounded in real depictions of stuff that is happening.
Again, what rare books? Like they pointed to a book with low published volume or foreign language is “rare”. If you have ever been to annual library book sales you start to realize just how many books get thrown out. I can go buy a stack of books from the 1880s and 1890s on eBay right now for $25. $4 each. Most people don’t realize how worthless books are, even “rare” ones.
I even imagine that the market price for these “rare” books is helping filter out anything truly valuable and rare. It just reads as a rage bait tmz article. The quantity of used books including “rare” books that get thrown into the dump is astronomical.
2 replies →
Something can be "rare" while also having no value. One only needs to take a look at their local Facebook Marketplace listings to see this in action.
Rare invokes images of limited edition runs of well loved books, when in reality it's probably extremely outdated software guides, how-tos, technical manuals, etc.
> Like Amazon consuming and presumably destroying rare books should be enraging to everyone, regardless of political persuasion.
No, destroying collectible books would be a shame, not just any rare worthless books. But these are not collectible. The article tried to dance around it by saying maybe some books have a sentimental value to someone somewhere. But that doesn’t mean any library or collector wants it. Don’t fall for manufactured outrage!
Edit: Here’s an example of an extremely rare book. My great great grandfather published a book of sermons around 1920. That book has zero value to anyone other than my dad. Would I be outraged if it ended up at someone’s estate sale, then a used bookstore, and then an LLM consumed it to learn to read? No; I would have expected it to have been discarded by humans before the LLM even got to it. Most of what we leave behind is discarded.
8 replies →
You don't know the books are being destroyed. There are automated scanners with page turners.
4 replies →
Should have let them torrent.
Just another reminder that piracy remains the absolute best archival strategy we have.
you do realize "rare books" in this context most likely means random technical manuals nobody cares about and not collectors items, right?
Oh, that's really interesting, I'd love to see the list of books that they've digitized too. Where did you find it?
follow the discussion around this on X, i think just a few minutes of research on this topic / reading past the headline you'll find out that this is pretty much a nothingburger
3 replies →
Unless we know this to be the case it's reasonable to assume it might be more.
It's reasonable to assume the clickbait article with no information is actually important?
1 reply →
Why would that be a reasonable assumption? This is so blown out of proportion. The imagery invoked by the narrative is one of huge corporations destroying the final copies of literary treasures. That's just not happening.
3 replies →
If most but not all, what number are ones that people and collectors do care about?
If collectors cares about them, the sellers wouldn’t have sold it for pennies.
2 replies →
Any bookseller worth their salt will ensure that a truly valuable book will not languish on their shelf for 1 minute longer than it takes to find a buyer willing to pay a fair price.
If AI-scanners are somehow bid-sniping bona fide collectors and wealthy aficionados, there may be cause for concern. But that is most certainly not happening here.
2 replies →
They are rare because nobody cares about them otherwise. Why is everyone acting as if they are trashing Gutenberg Bibles or first edition LOTR copies?
Carnegie was building libraries and concert halls, Silicon Valley CEOs buy girlfriends with big tits, cheat in computer games and destroy books.
Bad enough to steal IP but destroying books…that’s just evil
So is it moral of me to destroy a copy of "Windows 95 for Dummies"?
What if a person with small children and an elderly, incontinent pet with a penchant for peeing on books wants to buy it - can I sell this book to such a dangerous purchaser who might destroy it?
I love book as much as the next person but the hyperbole about "rare books" is absurd. Nobody is buying the Gutenburg Bible and destroying it for AI. The books in question are certainly not rare enough to be in museums - without titles there's no proof these are anything of real value.
To me, the hate against "destroying books" is kinda silly. The reason we're against destroying books is because the Nazis did it to destroy information.
This situation has absolutely no relation to that. This is the exact opposite, and the only reason this information isn't available publicly and is in risk of getting lost is copyright law.
TLDR: Amazon isn't the nazis in this story, copyright law is.
Making the verbatim content of the book lost forever is not equivalent to burning them? I'd even wager that the CO2 emissions of the scan,shred, and train process is even higher than burning.
1 reply →
But rich people want to violate copyright now! So judges ruled that it's now legal. But, because "the law is the same for everyone" they needed some excuse. NOT because, you know, obviously this demonstrates that very rich companies get to violate the law and you get hit by $30000 per infraction when you do it, when Anthropic ... doesn't even have to stop violating copyright when they do it.
Anyway ... this case is:
Bartz v. Anthropic PBC, No. 3:24-cv-05417-WHA, U.S. District Court for the Northern District of California, decided by Judge William Alsup
“the purchased print copy was destroyed and its digital replacement not redistributed, this was a fair use.”
https://copyrightalliance.org/wp-content/uploads/2025/06/Bar...
Obviously this violates precedent, I had an internal LLM (probably a frontend for Claude or ChatGPT) find them (just like the convictions for file sharing in the 2000s required counter-to-the-law reasoning by judges, fair use was almost never accepted as a valid excuse, even when it obviously was, but of course Sony was a billion dollar company and needed to be in the right. In fact that this had to happen was explicitly given as a reason to create the DMCA)
Anyway, some precedents:
Hotaling v. Church of Jesus Christ of Latter-Day Saints, 118 F.3d 199 (4th Cir. 1997)
“Although the Church acknowledges that its sole remaining copy is not the one it originally acquired … it maintains that the remaining copy does not infringe Hotaling's copyright because it is a replacement copy…”
(this reasoning was rejected by the court)
https://law.justia.com/cases/federal/appellate-courts/F3/118...
Atari, Inc. v. JS & A Group, Inc., 597 F. Supp. 5 (N.D. Ill. 1983)
... defendant sold a device for making backup copies of copyrighted Atari cartridges and argued that §117 permitted replacement/archival copying. The court rejected the broad replacement theory.
https://law.justia.com/cases/federal/district-courts/FSupp/5...
This very court has clearly declared that making a copy of a copyrighted work for replacement purposes is illegal, on multiple occasions.
I would like to point out that this isn't Anthropic's only extreme-WTF law violation. When the original judgement against them was made against them, they were forced to admit that using books to train models was illegal if acquired illegally AND THEN WERE ALLOWED TO KEEP DOING IT (thankfully the court never mentioned that part in the judgement so at least they can claim that was never decided when it becomes a huge problem in future cases as it obviously will). But that's not how this works. In my opinion Anthropic and OpenAI and everyone else need to at minimum take training material that was acquired in violation of copyright out of their training data unless and until they have a separate licensing agreement with the copyright holders. As long as Claude knows more about Harry Potter than is said in the promotional summaries it is obviously in violation.
Because of the copyright-filesharing court wars of the 2000s, which were also handled dishonestly by courts (whether we're talking US or EU courts), and the absurd copyright extensions, they had to now make some new excuse, and settled on this very sad, very destructive option. It's not even defensible legally, imho, but of course the biggest wallet must win. I don't understand. It's such a sad joke at this point, and it's not like the courts even still had credibility after the file sharing cases.
I wonder which sad excuse will be forthcoming from the courts when we have someone release a movie made by an AI model that is obviously a direct ripoff from some high-budget studio movie, and Disney needs to be protected from ... say ... "Scorched: Brothers of Aridelle" — In a vast desert kingdom, the royal brothers Elias and Anders grow up together, but Elias secretly possesses dangerous fire magic and isolates himself after accidentally hurting Anders as a child; years later, at Elias’s coronation, Anders announces his engagement to the seemingly charming Princess Hanna, provoking an argument that exposes Elias’s powers and sends him fleeing into the dunes, where he accidentally unleashes an endless heatwave that dries the kingdom’s wells and turns the capital into a furnace ...
I can think of a couple of solutions: 1. Glue a US $1 bill to the spine and take them to court when they destroy it. 2. Send them a license to use the book, with the condition that if they fail to return it within 30 days, they owe $1M. (Hey, if e-books can be licensed, why not physical books?)
In the end its not who can train the better model its who owns the better stack of training data. Acquiring old texts then destroying them makes their information proprietary.
How many AI companies are doing this? Say there are only 5 copies of a book left out there but 100 different companies want to shred it. Rare items could effectively vanish from the market. "Rare" meaning inclusive of high-quality items, disregard the bot opinion that mass book destruction is fine because most books are worthless anyway. And are any safeguards in place to make sure scanned books get saved in their original form for posterity before being recycled into trainingslop? Aside from Google's quasi-legal/ethical mass book scanning operation of course.
> "Rare" meaning inclusive of high-quality items,
This is the citation needed that is missing from every report so far, including this one which deliberately refuses to reveal anything about these books.
You’d think if there were examples of actually valuable, rare books being shredded that the journalists would at least be able to name one such example. Instead it’s always vague posting about the destruction without ever naming any examples.
I think it’s because if they named some example titles, everyone would see that they don’t care about these books being shredded.
On the contrary these companies should be publicly listing the name/info of every book they destroy and use as training data.
If the books are really worthless as you say then their case would be proved transparently.
This article presumably had to maintain confidentiality to protect the seller who agreed to place a tracking device in the shipment.
I would say the burden of proof is on the companies destroying human cultural heritage en masse, not the handful of journalists calling for attention to the matter.
1 reply →
I do work adjacent to the AI book scan-shred pipeline. There are definitely significant books that aren't "Windows 95 for Dummies" which are getting down to single-digit remaining copies.
I just looked for one novel, which wasn't a fantastic book, but it is the first use of a pithy and fun phrase that is so ubiquitous that you'll probably read it a couple of times today. I argued with Claude, GPT and Gemini for ten minutes just now, even knowing the title of the book, to even prove the book exists. It took me years to find a copy originally and then I lost it in a move. I found one more copy today from a rare book seller, but it just sold (to Amazon?).
Is the book valuable? Not particularly, but I feel it's noteworthy and important. I don't want to name it either, because now I have some searches out and the next copy that pops up I'll scan and put on IA. There can only have been a few thousand copies originally published in 1947, it's only in hardcover. I know of a couple of other copies in private hands, so it's not zero copies, but it has to be single-digits.
I have one periodical issue that I know of only one other existing copy (Worthpoint only shows one copy ever sold in their database) and if you look on collector sites there is a blank because nobody even knows what the cover looks like. I can't explain it, since the publication routinely printed hundreds of thousands of copies of each issue, but here we are. Perhaps all the copies were withdrawn and pulped immediately after publication for some reason? It's in my scan pile, so I'll have it uploaded soon. Is it significant? Not hugely, but every other issue of this title has been scanned already, so it's scratching an itch to get this one done.
There's definitely rare stuff getting scanned and shredded. Someone in a comment above said it's not like Nazi book-burning since they were trying to destroy information. But it is like that if you consider there remain no other physical copies and all the electronic copies are locked up in a way that nobody can access except to trick an LLM to spit out a paraphrased copy from its training data.
TechCrunch is trying too hard to make a connection in the title there. Amazon was trying to make money then as it is now. Their selling of books at the beginning was no more principled than the destruction of books now.
This is the legal loophole that allows them to do what they need to do, and the benefit of doing it outweighs the modicum of outrage this title will generate.
Edit. Because I see my statement confused all the HN experts:
Anthropic's version of this was Project Panama [1]. The destruction of the original allows them to keep a digital copy in the AI training dataset. Quoting from the page:
> Judge William Alsup ruled that the destruction and digitization of legally purchased books constituted fair use
[1] https://en.wikipedia.org/wiki/Project_Panama
Is it a “loophole” that you can buy something and burn it? That’s how property ownership works. If the books were really so rare and valuable then the original owner should have given them to a museum rather than sell them as scrap. Yet no one who is complaining right now cared about them before Amazon got involved.
> Is it a “loophole” that you can buy something and burn it? That’s how property ownership works.
It sounds shocking if you're missing the context and only rely on the blogspam which is this TechCrunch article.
They aren't buying the books only to destroy them but to build an AI training dataset, which means they're making a digital copy. This becomes a copyright issue. The destruction is the legal loophole that allows them to keep the digital copy as the only copy in circulation and be considered fair use as decided by a judge (see below).
The concept of ownership means you can do anything you want with the object, the book in this case. Not with the content, like make or distribute copies. The scanning machines are cutting the spine of the book, feeding the pages to scanners, and then destroying the physical copy.
Anthropic's version of this was Project Panama [1]. Quoting from the page:
> Judge William Alsup ruled that the destruction and digitization of legally purchased books constituted fair use
[1] https://en.wikipedia.org/wiki/Project_Panama
No, that's not the loophole. The loophole is that a US federal judge recently ruled that scanning a book for the purpose of training an AI is "fair use" and hence legal _if and only if_ the book is destroyed in the process. If you retain the original, you would be creating an unauthorized copy, but if the original is destroyed, then it is permissable as format shift: https://fingfx.thomsonreuters.com/gfx/legaldocs/jnvwbgqlzpw/...
Sometimes original owner does not know that it is a first edition single copy. For them it is an old book. There have been many cases where owners did not know true value of their antiques.
1 reply →
Its somewhat ironic.
The worst part of Amazon destroying priceless old books in the quest to build the torment nexus is the hypocrisy
They obviously aren’t priceless if Amazon is buying them for pennies.
all due respect but if they buy them they can do whatever they want with them. not like they are people and not like you have a right to them either. not even like you would have been able to acquire them yourself.
2 replies →
> As the bookseller who sold them told me, there are not many people in the world who would care about them in the same way people might care about the first edition of Oliver Twist, but that doesn’t mean they’re not valuable.
So books that would probably have ended up as trash. These AI training facilities are actually doing these book a service. Not only they are probably going to keep the scans safe (for future training), but having the book end up in an AI model may be the only way it is going to have any use at all.
Let's say for instance that the book in question is about woodworking, and it is not great, lots of mistakes and inaccuracies, unoriginal content, etc... except for a single thing, maybe a trick for making a certain measurement or something like that. Who would read such a crappy book for this single good trick he doesn't know is there, well, an computer will, computers process terabytes of crap without tiring and complaining, that's what they are for, and with a well designed LLM, that one trick may resurface, waiting for someone to ask about that specific measurement.
The problem here is not that rare books end up in AI training facilities, it is that these AI facilities are owned by for-profit companies keeping the data to themselves. These books should go to public libraries instead, for everyone to access, it would be better if these books weren't destroyed in to process too. But the question becomes: why didn't public libraries didn't do that in the fist place? And maybe in a more respectful way. The AI companies would just have had to license the database to libraries, probably simpler and cheaper than having their own scanning facilities.
To me, this mess is a failure of the copyright system. One one hand, large scale book digitization projects intended to preserve and make the original text accessible get lawsuits by publishers, while AI training is "fair use". It means we have built a system that encourages destroying rather than preserving these books!
Yes, we've inadvertently set up a system of incentives that were designed to preserve information and monetize it. And we've instead set up a set of incentives to make it scarce and destroy it.