AI companies are shredding rare books

5 hours ago (twitter.com)

https://xcancel.com/HedgieMarkets/status/2081534588485296565

I've limited sympathy for the publishers.

It pisses me off to reflect that they can sit on works until copyright expires, keeping them out of print. There's no real need for any of these so-called rare books to be rare while they're under copyright.

And related to this, the books that are in print are mostly only in print in the shittiest way. I often see well-made books from the 17th or 18th centuries which are still in good nick. It's ridiculous that in the 21st century, publication standards have fallen to the point where for most works a disposable format is the only type available - where no amount of money could buy a truly decent hardback copy.

If we have to have copyright laws, I'd like to see two changes to them.

When a publisher has no incentive to keep an edition in print, it should be available to any other publisher to print, without compensation to the original publisher, and with renegotiated royalties for the author.

And if the publisher keeps a book in print - but only in bestseller-grade materials, bogroll paper that furrows in any humidity and perfect binding that molts its pages a couple of dry seasons later - and if it refuses to print a durable hardback copy with signatures, good paper and decent print - something that will still be readable in several generations' time - any other publisher keen to have a crack at it should be able to, again without any compensation for the original publisher, though perhaps in this case, with matching royalties for the author.

  • A friend proposed that copyright should just die with the author and / or their spouse and I'm left agreeing. I want books and music to be less strict on copyright. Some of my favorite YouTube channels break down music and songs, and go as far as recreating beats / tracks from famous hip hop songs, but someone at a record label company dings every one of their videos, they can barely sample a few seconds, its VERY CLEARLY fair use, and even gets me to listen to the songs more, but they are stealing all revenue for a fair use video that took a lot of time and money to make to begin with, they are profiting from work they didn't even do. This is silly to me, I'm sure there's loads of other channels and videos out there screwed over by the record industry over 15 second samples from a song... Which should 100% be fair use and not be given to the record labels.

    Issue is small streamers have no legal defense.

    Edit, adding the channel I am referring to:

    https://www.youtube.com/@diggingthegreats/videos

    • > copyright should just die with the author

      That would have a few undesirable consequences... for example, you wouldn't hire a 70 year old writer for your commercial project no matter how brilliant they are.

      The complexity of our legal system is in many cases justified. The problems are often the numbers (duration of copyright protection etc.)

      8 replies →

    • 20 years to make some money, and then we set the work free for the public benefit. If it's good enough for patents, I don't see why it isn't good enough for copyrights.

      7 replies →

  • Books in the 17th and 18th century often didnt get bound by the publisher. They were done by independent binders for _custom_ orders. You’d see whole libraries with the owners binding/cover standards rather than per book.

    Books in that time were _luxury_ goods. Most people could not afford them. One of the ways that was changed was to introduce cheap, mass produced bindings that were lower quality than the bespoke artisianal bindings done by specialist craftsmen.

    You can still get custom bindings done. There exists whole niches on the internet of crafters that will take a production run book and strip its binding and make you extremely high quality and custom bindings and covers.

    • Learning this fact is what got me interested into binding my own books!

      With how the quality of things seems to have been degrading over the years (either real or just me getting older and experiencing the impermanence of all things) I've been trying to adopt an attitude of "if this practice existed before the industrial revolution, I can _probably_ do it" and it's been really great to learn how things were made before they had to be mass produced as cheaply as possible.

    • One of my relatives is one of those micro publishers. Selected works are printed and bound to extremely high standards and materials in a way that their customers are willing to pay 1k-25k+ per book. These editions only have a couple of prints and are mostly made to customer requests.

      There is a market for these type of books, albeit a very small one.

    • You don't even need custom bindings —- library bindings are a thing, even if they cost significantly more then a typical hard cover binding.

    • The Czech National libraries do that for newspapers, magazines and other periodicals - they bind the copies they get automatically for preservation to big books, so they can be better stored in their archives.

      Apparently it is getting harder to find people who can do that as most schools no longer have book binding as a course you can study.

  • I have even less sympathy for IP stealing LLM operators

  • > I often see well-made books from the 17th or 18th centuries which are still in good nick.

    Those books predate the development of wood pulp paper. It isn't the publisher's fault they can't economically print on rag paper anymore.

  • > It's ridiculous that in the 21st century, publication standards have fallen to the point where for most works a disposable format is the only type available - where no amount of money could buy a truly decent hardback copy.

    This is purely a response to market demand. Publishers aren’t going to put in the extra expense of binding high quality versions of every book so it can occupy warehouse space while consumers everywhere buy the cheap paperback.

    > And if the publisher keeps a book in print - but only in bestseller-grade materials, bogroll paper that furrows in any humidity and perfect binding that molts its pages a couple of dry seasons later - and if it refuses to print a durable hardback copy with signatures, good paper and decent print - something that will still be readable in several generations' time - any other publisher keen to have a crack at it should be able to, again without any compensation for the original publisher, though perhaps in this case, with matching royalties for the author.

    This is all based on the idea that there is hidden demand for something, but publishers are choosing to deprive us all of it for reasons. That if we open up the laws, another company will come along and satisfy this hidden market opportunity and associated profits that publishers are declining to take.

    The simpler explanation is that these high quality editions aren’t being published because the publishers have the data about demand for them. They know they won’t sell.

    If the goal is preservation, laws forcing publishers to print on slightly nicer paper isn’t going to solve the problem. It needs to be a robust digital archive and it needs to exist somewhere other than in unsold warehouse inventory or some book collector’s shelf. You’re trying to solve a problem with last century’s technology.

  • I'd go for the less complex version: just make copyright last for like 20 years at most.

  • > I've limited sympathy for the publishers.

    This doesn't affect publishers.

    It affects humanity as a whole by having large private parties hoard books to partly destroy them. The covers and spines can have historical value too. There's not even a reason for these companies to release their scans to the public once the copyright expires.

    It prevents proper preservation by archivists and preservationists.

  • No one reasonable is concerned for publishers they are concerned for the books. Assuming the allegations are correct historical artifacts are being destroyed with no trace. It doesn’t even matter if they republish it. The artifact itself and the time it has transited is the valuable thing. It is simply immoral to destroy stuff like this.

  • Yeah it's as if the publishers are somehow strapped for cash compared to AI companies.

  • Thanks for your thoughts, but how is this is anyway relevant to the issue at hand? I'm sad that this is the top comment...

  • I've never read a coherent argument for why all of these "should"s should be. I understand you want to read the books. But I don't see how that desire results in the laws needing to be changed so that you can read everything you want, even things the author or owner doesn't want you to read.

    There is not a shortage of books in print. I don't understand why people get so hung up on a few of them being out of print. It is okay for someone to own something cool and not let anyone see it. The cool thing doesn't suddenly become a societal necessity because it is words written down.

    • > There is not a shortage of books in print. I don't understand why people get so hung up on a few of them being out of print.

      Because the point isn't so much in reading whatever text as if all text is the same. The point is the spreading of knowledge. A single book can contain knowledge not present in any other.

      > But I don't see how that desire results in the laws needing to be changed so that you can read everything you want

      In order for your "people" (country, etc.) to do better, you want them to be educated. In order for people to understand one another, you want them to be able to see all the same various perspectives there are. It makes perfect sense for laws to aim for these goals. This is why libraries exist.

      > It is okay for someone to own something cool and not let anyone see it. The cool thing doesn't suddenly become a societal necessity because it is words written down.

      Books aren't simply trinkets, like an item you bought at a gift-shop.

      > I understand you want to read the books.

      I think you're looking at this too much as what people want for their own individual selves, when it's more of what people want for everyone. It's about what they believe is best for society as a whole. They don't need to want to read a book themselves.

      1 reply →

  • > When a publisher has no incentive to keep an edition in print, it should be available to any other publisher to print

    How do you adjudicate that? And wouldn't it just lead to loopholes such as "Ghost Printings" (cf. Ghost Flights https://en.wikipedia.org/wiki/Ghost_flight_(commercial_aviat... ) where the books are technically printed in the required volume but practically unavailable to customers through one method or another. Because, the cost of wastefully printing a few books to warehouse, is less than the potential losses of the IP rights, probably.

    • I think putting the works in a print-on-demand catalogue at an unreasonable price would be more likely.

      But there are worse outcomes.

We reprint old books after checking out copyrights (for all books, this means pre-1930, but for some (I'd actually say most) it also means ones published up to 1964 and 1973, depending on how the rightsholders did (or didn't) do the renewals).

We use a special guillotine type cutter to cut off the binding and then store the pages in a sealed plastic bag which goes in the archives; they're stored there indefinitely in case the book needs rescanned for some reason. We also keep the original, uncompressed copies of the books on magnetic disks.

We also go out of our way to try to find rare books published in 1931, 1932, etc. so they are ready to go once the copyright expires.

And no, no AI company has ever come to us and asked to run training on all of our scanned copies.

  • > We reprint old books after checking out copyrights

    Who is we?

    > then store the pages in a sealed plastic bag which goes in the archives; they're stored there indefinitely

    Is that the best thing for archival storage? Like could things like chemical breakdown increase the humidity in the sealed bag or concentrate corrosive chemical vapors? I was under the impression the best environment was an actively climate-controlled environment.

    • The issue is mold. Sealed plastic bags for old books is a recipe for mold. At least they're individually stored, so they would develop mold at independent rates. (And maybe the seal allows airflow.)

      Most books have some mold by the time they're ~100 years old, it's just not enough to cause a serious problem. Sealed enclosures (wrappings, bags, tubs) are a nightmare situation. Even packing books too tighly on shelves accelerates mold growth to problematic levels.

      3 replies →

  • Some source or citation or context is needed here, is this the work of a 2 person no profit or a trillion valued pre IPO company?

  • > And no, no AI company has ever come to us and asked to run training on all of our scanned copies

    The value of most very old books for AI training is very low. You don’t really want your AI training data to start biasing toward outdated writing styles. Most of the valuable knowledge has been covered again in modern texts in more depth and detail.

    There is interesting value in old texts and it’s important to have them archived. It’s less valuable for stirring into the giant pot of AI training data, though.

    • If you only care about facts, maybe. Even then I'm sure there are countless facts not described outside of old books.

      I have a hard time believing that text valuable to humans would not be valuable to AI.

      1 reply →

  • > And no, no AI company has ever come to us and asked to run training on all of our scanned copies

    Yet

  • Why cut off the spines? Isn't that how you end up getting unattributed 'dead sea scrolls'?

    • It's easier to get good quality scans from individual pages than it is from pages in a complete book. Imagine laying a book flat, then the page is distorted in the area around the spine. You can work around this (either by trying to correct for the distortion in software or with clever scanners that position the book more advantageously) but it adds complexity compared to chopping off the spine and just dealing with flat sheets of paper.

    • Easier to send them through a duplex scanner (or put them in a flatbed one if they are fragile). Cheaper than buying the automated ones with the page turning robot arm.

      I have done it at home for my books since the mid-00s.

      3 replies →

  • or just send the books/scans to a country that doesn't recognize US copyright

> You can reprint a bestseller. You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal. So it's going to accelerate.

Aren't they shredding only the books still under copyright protection? How is an 18th century botanical text still under copyright?

IDK about the shredding, it's not nice, but it's more a problem with copyright law than AI companies.

Scanning books you own should be legal from a copyright point of view, and not require shredding.

Second, one should think about abandoned property provisions for copyright works published more than 50 years ago and in danger of being forgotten: once challenged, either you as the owner have to prove that the work is preserved for future generations (e.g. in various libraries around the world), or you have to authorize further copies, or you give up copyright on the work.

  • Scanning books by taking them apart into singular pages and scanning those pages is faster and cheaper. AI training is a numbers game, so they want faster and cheaper.

    What happens to the pages after? No one needs them anymore, so they get mulched and recycled.

    That would be the dominant scanning method even if copyright wasn't a thing. But then again - if copyright wasn't a thing, there would be much less need to scan any physical media.

    The reason why OpenAI can't just go on Amazon, buy a "digital edition" of a 2018 book and use that is that it would violate the license in ten ways, and then the DMCA laws that forbid breaking DRM on top of it.

    • Think about that last point for a moment. Our “rights to read” are diminished significantly with digital works as compared to printed works. Right of resale. Right to lend.

      In the end, digital publishing just isn’t right and will lead to massive gap in our historical records. They require active curation and cannot be preserved simply by resting on a dusty shelf.

      2 replies →

    • > What happens to the pages after? No one needs them anymore, so they get mulched and recycled.

      Strictly speaking, no one needs the Sistine Chapel or the Pietà etc. It would be a shame if they were mulched and recycled, though.

      1 reply →

    • Machines for non-destructively scanning books were developed and perfected long ago. The destructive scanning is neither technological limitation nor an issue of expedience. It's an issue of copyright law and fair use.

    • > if copyright wasn't a thing, there would be much less need to scan any physical media.

      Because there’d be much less content created in any media to capture in the first place.

      1 reply →

  • > You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal.

    An 18th century book would be out of copyright so why would it be illegal to keep the original and scan it?

  • > Aren't they shredding only the books still under copyright protection? How is an 18th century botanical text still under copyright?

    It's cheaper to scan the books if you do it destructively. Cost. That's why they're shredding irreplaceable texts. Nothing to do with copyright.

    https://www.404media.co/ai-companies-are-buying-tons-of-old-...

  • Old rare books where there are single digit copies should enjoy some sort of patrimonial protection just like museum pieces. You can own them but have the state have the option to buy it if you’re about to significantly deface it or destroy it.

  • Books that are shredded can’t be scanned by competitors.

    • Yeah, this is the point. I don't understand the bulk of this conversation. Copyright doesn't matter, the books themselves don't matter. All that matters is that their corpus of training data grows faster than their competitors.

That's why archive.org should have never been sued for lending books they had physical copy of. This is the result. Publishers should be more careful what they wish for.

  • The archive.org story was more nuanced than that. If I recall correctly the full story was that they used to lend digital versions of books they physically bought and scanned with DRM to enforce a sort of one to one at a time restriction.

    But during covid archive.org decided to just remove the limit and lend unlimited copies concurrently which started the debacle with the publishers.

    • > But during covid archive.org decided to just remove the limit and lend unlimited copies concurrently

      IIRC, this was 100% it. Lending one digital version of one physical asset was likely already a violation copyright. Lending UNLIMITED digital versions of one physical copy was DEFINITELY a blatant violation of copyright.

      5 replies →

    • Yeah that was it; if I got this right, US libraries got the right to lend out one digital version of a book that they had in their inventory. Archive.org combined those digital versions so that people could check out a digital book if any library in the US had it (digitally) available. But during the 'rona they removed this limit and just lent out books regardless of it being "checked out" digitally from a library.

      This wasn't a very smart move of them. I get why they did it but they put themselves at a huge legal risk.

    • If i recall correctly, they had permission from physical libraries to use their copies as well, so it wasn't just a single copy but many copies, just one of them was converted into digital. Still wasn't enough apparently...

  • Publishers don't care if rare books get shredded?

    • And, regrettably, The Archive lent books regardless of physical possession.

      Publishers had accepted the prior arrangement before The Archive decided to push it, if not explicitly then implicitly by not suing.

      I'm a believer in The Archive's mission, and I wish they had treated the goodwill they'd accumulated as something worth preserving and not a currency to be spent.

      It has been stated by many before me: lending books should have been handled by a separate entity, especially when they removed the physical backing requirement.

      2 replies →

    • Yeah, why would it be bad for publishers? If anything they'd most likely encourage more book shredding!

    • There are many reasons but it's just rare books - in general: they are setting a precedent for people to pirate instead of using libraries. In general companies are trying really hard to make people switch back to torrenting, pirate sites, sharing media etc.

  • > That's why archive.org should have never been sued for lending books they had physical copy of. This is the result

    How are these things remotely related? If anything, Archive.org’s callous, thoughtless approach nuked the hands of legitimate archival efforts.

The publishers sued AI companies for training on shadow library data, hoping to negotiate content deals for big $$$ down the line. Instead, they got analog hole'd.

Turns out that buying an old book for $5 and destructively scanning it for $25 is way cheaper than paying extortion fees to the copyright-mongers.

What I don't buy is it being "rare, precious books". First, they're not after ancient texts - they're after the books that there's still copyright on. Second, when it comes to books, "old" doesn't mean "valuable" - plenty of libraries destroy old books because there's no demand for them, and storage costs you. This is how those scanning companies get books for so cheap.

  • This is such propaganda lol

    They change the info (many buyers), destroy the source material, and now the lie is in the LLM.

    That's it!

  • So they're not valuable... except to AI companies. They should pay a fair amount.

    • They are paying a fair amount. In the ballpark of $5 per book.

      You know, piracy online is nice and simple - but it's kind of hard to get physical media without paying what the previous owner considers "a fair amount" to part with it.

    • > They should pay a fair amount.

      they should pay the marginal value that the next buyer would buy.

      Do you also think that a person dying of thirst ought to pay the maximum price they could possibly pay for water?

      8 replies →

    • Why are you assuming they aren't? The people who have those books are choosing to sell them for a price. What isn't fair?

    • They're paying the market price for a used book and getting the same rights out of that purchase as you or I would if we bought a used book.

  • I see.. so the various AI companies are in the right on this?

    • I think they are in the legal sense of right, and I think they only discarded the remains of these dissected books because previous rulings (e.g. archive.org's lending practices of digital copies of books they physically owned) gave rise to a situation where destruction bore less legal risk. As for the moral case, I don't have much to say on that, we all have our own lines in that sand.

      3 replies →

    • yes. in the sense that they probably wouldnt want to do this but the law forces them to do an extremely dumb thing.

      they're paying for the books, no shady things going on there. whether the publishers should deserve more than a single copy's worth is a separate question.

      having the law such that its illegal to scan a book and then keep it, but legal to scan it and destroy it - gg no re there, law people retardmaxxed themselves as they tend to do with anything related to digital data.

      2 replies →

  • > Turns out that buying an old book for $5 and destructively scanning it for $25

    More like destructively scanning it for $0.25, if you include the wear and tear on the machine, the salary of the guy who unjams it when it chokes, and recycling fees.

  • Buying up old books is legal. Digitizing and distilling them arguably is too. And also:

    > paying extortion fees to the copyright-mongers

    Yes yes, greedy fat cat publishing oligarchs treading on the poor put-upon scrappy AI underdogs. /s

    Back in reality, the fraction of people who got into publishing books to get rich collecting rents is… not large. There’s so many other fields that are likely to reward participants with more wealth that it’s absurd — even with all the passion for the work in tech it’s probably relatively less pure.

    And whatever the excesses of copyright have been, the whole bargain has always been on more pro-social foundations and stronger intellectual foundations than “extortion” sneers. It recognizes that incentives matter and work that’s valuable should be rewarded and incentivized.

    A culture that takes a Robin Hood approach to low marginal cost billing points but fawns over the hypercapitalized distribution King Johns isn’t creating a freer or richer society or fighting the real cartel center, it’s indulging resentment and caricature.

    • > Back in reality, the fraction of people who got into publishing books to get rich collecting rents is… not large.

      That's because the business is buying copyrights in bulk from those people. Copyright-mongers ≠ publishers or writers.

I see mentions of Bradbury's "Fahrenheit 451" in that thread but what this really seems to be mostly like is Vernor Vinge's "shred and scan" factory in his novel "Rainbows End".

  • Perhaps the last great near-term predictor. I often wish he'd written more. To remind us all, he predicted shred and scan would be a short stop over done by villains on the way to nondestructive scanning.

    That said, supporting Anna's archive is one of the best things you could do for humanity long term in my opinion.

  • Well it's close to the author's intention for Fahrenheit 451, but just not what everyone wants it to mean.

Maybe we can kill two birds with one stone: digitize rare books and reverse the damage from Authors Guild v. Google [1].

Let AI companies do this. But require them to make the digital copies public. Maybe with a multi-year delay, to give the original scanner advantage to doing it.

[1] https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,....

  • Yeah, I was thinking along these lines...

    Let's say one of the books to be digitized and destroyed is the sole remaining copy of a book from 1850, which is now considered public domain.

    On one hand, hoarding such a book, stealing its content from the public domain, locking its content behind a for-profit machine, and destroying the only remaining copy is clearly wrong. It's equivalent to stealing a public resource, just like mining minerals or oil on public lands without a permit or mineral rights. Pure extraction.

    On the other hand, taking care to digitize the copy and making it available for free in perpetuity, as well as being required through regulation to provide access to that content through, let's say a public utility LLM/AI available for free through libraries and online... and perhaps after fair due diligence being required to preserve physical copies in a public archive of rare books of which there are no known remaining physical copies...

    That seems much more reasonable to me at least. I can imagine there are many who would not see it that way though. Do we see it happening or gaining regulatory, moral and/or public support?

    • I think this misconstrus what public domain is.

      It provides a freedom to circulate, but not access to the material. It is not a public owned resource.

      Turning a copy over to the public or state might be an interesting requirement for obtaining a copyright, but instituting that fix for new works now would have a 70 year lag time.

      Think of it this way, if I copyright a book and put it in my dresser for 70 years, that doesn't give the public the right to access it or come into my house and scan it after expiry

      3 replies →

It seems a key contention of theirs is the possibility that rare books are being destroyed this way, yet the things they cite don't seem to suggest this (based on their paraphrasing), they just throw the following at the end to make it seem like it's occurring to irreplaceable books:

> You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal.

Is there evidence of this? Since otherwise they could very well be describing what is only occurring to in-print or non-rare books. (This is a genuine question since their post doesn't shed any light on it.)

  • There is zero reason to shred 18th century books. Any such books are out of copyright.

    • Yes, that is what the person you are responding to is saying. That's why they are questioning the assumption that this practice extends to 18th-century books that are solidly in the public domain.

      If it is happening, it is an outrage. However, the 404 article doesn't actually provide any evidence of this; it just connects the shredding of digitized books and the digitizing of rare books to the assumed shredding of rare books, which isn't necessarily happening.

From what I understand, the rare books in question are not some historically relevant medieval manuscripts, but rather some relatively recent books (still under copyright) for which there are few print copies available for purchase.

I am not sure that physically destructing one copy of this type of book to preserve its contents digitally is so bad. Pretty much anything that is still under copyright should be valuable only for its content, not for the physical medium it's printed on.

  • > historically relevant medieval manuscripts,

    These are increasingly available online, btw. Historical research is accelerated when historians have direct access to scans of relevant source material. Not destructively scanned, of course.

  • > Pretty much anything that is still under copyright should be valuable only for its content, not for the physical medium it's printed on.

    Well, there are special collectors editions with the signature of the author and gold pages or whatnot. But the AI companies are probably not using those.

  • Is slightly modifying some weights in a markov chain somewhere really preserving it's contents?

    • It's not really Markov chain as you need full P(x_t|x_{t-1}, x_{t-2}... x_1) instead of just P(x_t|x_{t-1}).

    • They used to put them all up on books.google.com until they were ordered not to if they were less that 100 years old. Most of the stuff that was accessible was copied to archive.org, and almost all of that stuff ended up on annas-archive and library genesis.

      Consequently, many/most books are more available now than they've ever been. This is mostly a copyright question, not a question of preservation. I say mostly, because those scans, without redundancy and fingerprinting, can be changed and bowdlerized in the future without remaining physical copies as a reference.

      I own about 3500 print books, and started a project to find scans for all of them (that I need to get back to.) Average publication year is probably around 1975, and the bulk ranging from the 1940s to the 2000s. I made it through about 1500, and couldn't find maybe 40, most of them bad. e.g. self-published stuff like "My God Heals, Does Yours?" This was a few years ago, if I went through those 40 now, I bet I'd find half of them.

      I hate what the AI companies are getting away with, but only because they get to violate copyright while being aggressive enforcers of copyright and DRM circumvention laws. Destructive scanning, however, is a cheap way to get a good copy of a book online. If that book were then put in a place where teveryone interested in its contents (or who are just hoarders) can get a hold of it, you'll be able to find a verifiable copy of it 1000 years from now.

      As of now, Russia and annas-archive are just a few points of failure that can erase those words forever. It will be done with armed, uniformed men, and people will claim it is not dystopia, but justice. I don't want the only copy of the text of a book to lie in the interpretation of some privately trained LLM.

Reminds me of Blood Meridian where The Judge meticulously sketches the rock glyphs that he comes across, and then destroys the original.

Now that I think about it, The Judge is an apt metaphor for AI : "Whatever in creation exists without my knowledge exists without my consent."

  • No, it isn't. Nothing about training a model requires this. You're thinking of copyright, which does treat information this way.

    • Sure, I'm conflating the tool w/ the people building the tool - but imo muddling the distinction is good, because the hard distinction is what tricks people into an unthinking mode.

      2 replies →

    • Well, AI and Copyright are both related things that people created, so it's certainly relevant. I don't think the person you're responding to misspoke.

Isn't this the legal requirement for digitizing books usually? I dont like that its being done, but I feel the direction of anger for this one is misplaced. Or at least partially misplaced.

  • Yes part of the legal defense used to good effect is that they are copying one for one (not really making a copy but rather converting it). I agree with many in this thread that copyright reform is the best solution here. Shorter copyrights. Lose copyright if you stop using it (printing it), more clear digitization and preservation guidelines.

It's just modern book burning. The contents don't even matter, you will never see the text and images in these books again.

They should be forced to publicly release the books as an Ebook.

Leave it up for anyone to download and then compensate the copyright holders later.

In fact if ingesting these books for LLMs is fair use, us commoners should be able to read them for free. Maybe restrict commercial redistribution though.

  • There's a bunch of details regarding LLMs and copyright, but I don't see how

    > They should be forced to publicly release the books as an Ebook.

    would be reasonable in any way? The books aren't theirs to release publicly. If I brought a copy of any given movie on DVD, ripped it and used it to train my own "LLM" located at /dev/null, should I then be allowed (or even forced) to release the movie publicly for anyone to watch for free?

    • We’re in a brave new world now. If theirs reason to believe your destroying the last copy of a book( or another copy isn’t easy to obtain) then you should make a copy of it available.

      Set up a compensation fund for the rights holders. Anything is better than culture literally being sucked into the void.

I dunno.

It's sad in a romantic kinda way, because of the lost artifact, but the information is what makes the book valuable, not really the medium.

The out of copyright books don't really need to be destroyed anyway for them to be fair use for AI training, and arguable, even if you needed to, you only need one copy per title per company at most.

So it's not a gigantic loss.

  • But is the information really scanned and preserved for direct access? Or is it just trained into an LLM so that we can only get fuzzy answers about the information?

>And the judge said it's legal.

Was was the alternative? Order them to stop quietly destroying their property, it is making someone else very upset?

Which book that was rare was destroyed? I'm interested to know a few titles.

  • The 404media article mentions notably https://nltimes.nl/2026/06/25/rare-book-dealers-fear-tech-fi... which says:

    > The attachment contained 3,000 English-language titles organized by ISBN number, including books such as Distinct Element Modelling in Geomechanics by K.R. Saxena (1999); Barrett's Traditional Fairy Tales (2021), an academic study of Irish folklore; and Laser Shock Peening of Advanced Ceramics by Pratik Shukla (2018).

  • Rare or not, destroying books is in itself a morally repugnant act. I don't know man, it's not so long ago we used to view nazi book burnings as an archetype of evil. Today companies are offering book-burning-as-a-service and it hardly causes a stir.

    Bearded German man was right.

Librarians were already doing this at scale in a process euphemistically called "weeding":

https://www.ala.org/tools/challengesupport/selectionpolicyto...

  • This is a bizarre comparison.

    Weeding is the natural process of disposing of less-demand books. Like the rest of us, libraries operate in finite space, so if they want new books, they have to remove ones their users aren't using. Most libraries will try to sell books before disposing of them in any destructive way.

    What similarities do you see here?

    • People are very desperate to try to claim that something that's somewhat obviously morally wrong is actually highly nuanced, because it makes them feel uncomfortable.

      4 replies →

  • Libraries are not archives, and weeding does not mean destruction.

    • Libraries in my area dispose of 'weeded' books by covert means, so that public ire isn't aroused by finding dumpsters full of discarded books.

      They are most assuredly destroyed.

  • This is really dishonest framing, unless you really, honestly can't tell a difference between pulping a mass market paperback romance novel that there's 3 million of in circulation, and shredding an 18th century botanical text that there's only 2 copies of in existence.

    • You call of dishonest framing, but you're begging the question twice.

      Once that AI companies are really shredding 200 year old rare books, and once that libraries are only weeding mass market pulp fiction.

      4 replies →

I had some old computer/unix/etc books for which I couldn't find digital copies. I was considering paying to have them scanned because lugging them around each time I moved was getting to be a pain in the butt. When I found out they destroyed the book in the process I could just never go through with it. I later found out there are non-destructive scanning machines (even open source ones!) but ended up selling the books before I ever went that route.

How much of this shredding them isn’t just copyright but rather they don’t want anyone else having this information in their datasets?

Horrendous stewardship of humanities collective knowledge all for profit and the race to have the one god computer to rule them all.

As more time passes it becomes clearer that America’s AI strategy should’ve been a public private partnership where the public owned the datasets and the underlying models and we’d leave the productionizing of LLMs to private businesses

  • Honestly, I think its just the fastest way to scan them (there are slower non destructive methods available too!)

    Also, they can then just recycle/dispose of the paper and don't have to worry about reselling/donating the books themselves. I suspect this is all about speed of data ingestion and anything else is a side effect they don't care about.

You can’t champion on the greed that is copyright and then be sour when Anthropic tries to work within this framework.

There is a word for that kind of behavior.

This is the result of our copyright law in the United States, which is extremely tilted to favor authors and publishers. The judge made exactly the right call and the companies are following the law.

The fix here is to change the law to permit training AI without destroying the original materials. But that is going to be a heavy lift.

This is the opposite of a book burning. These books which only a few would ever know the names of, let alone find, let alone read, are being digitised so they can be found in electronic searches.

  • > being digitised so they can be found in electronic searches

    You make it sound like they are running a second Project Gutenberg. They most definitely are not making these available for electronic searches. At least not searches the public can participate in.

    • Correct. If they were doing this with Gutenberg/Smithsonian/some library so there's both a public archive and training the language model, it wouldn't have the ick factor.

  • And knowledge from these books can be "imprinted" into a model to be used by much more people or even survive this planet once sent into space in a probe.

  • > are being digitised so they can be found in electronic searches

    What guarantee do we have that the book contents will be served unfiltered and unaltered?

  • Are they being digitised? That word implies that the book would have a digital representation of the original, which is not what an LLM is. Can they be found in electronic searches? I mean I kinda get what you mean, the knowledge was potentially forgotten and now it might not be, but also the authors who created that knowledge get no credit, no reward.

Digitizing books, even if it means the destruction of the original, means more people end up having access to the knowledge within, and is a good thing. Full stop.

  • They are digitized for private training and not for public use. Not a single human is reading these.

At the very least why not upload the scanned books to the internet archive while already at it?

This is highly disturbing news; is this standard practice? What did Google Books do before?

  • Yes, this is very disturbing to me. I recently discovered how great old, out of circulation/print can be. Now old abandoned libraries are like a treasure trove to me. If these books are not digitaly preserved, that sounds very sad to me.

Here's the original story from 404 Media that this tweet is paraphrasing (without linking to, of course, because X disincentivizes links): https://www.404media.co/ai-companies-are-buying-tons-of-old-...

The segment that talks about rare books:

> One professional bookseller who specializes in selling foreign language books on these marketplaces told me that, starting in April, he and other booksellers noticed a historic spike in sales. [...]

> This bookseller said his inventory is full of rare, foreign language, and low circulation books, meaning that if they are destroyed in the process of becoming training data, they’ll be even harder to obtain.

Literally Rainbows End. Vinge was an oracle.

It's a real shame that no one ever got that book in front of Hayao Miyazaki's eyes.

  • It's amazing what reading a word in that book like sousveillance can do.

    I looked it up.

    surveillance, sousveillance - sure, why not? Seems harmless.

    But it's as if the more people know the word sousveillance, the word and even the act of surveillance actually loses a little bit of power (its embedding changes, if you will!)

    All our fears about the end state of surveillance can now be countered by an end state of sousveillance.

    Before I had sousveillance to think of, I could only think of surveillance (when thinking of veillances) - and it was more of a threat then than it is now.

    The world is made of language, or as Terence McKenna said made out of words which sounds obviously false at first. But go looking for the inside of an atom and tell me what you find, and think about where the medium of reality actually implements itself.

  • Well, an oracle except for that part in the same book where unlimited access to the internet made all children prodigies self-motivated by their love of learning.

Could someone explain to me, why rare books are so precious for them?

Say i feed the largest LLM a book of an alien civilisation, that it definitely hasn't seen before. Then this tiny piece of text muds the vast ocean (latent space) of the model minimally. It will not be able to cite from that book reliably after that fine tuning. Especially for rare books, because them being rare implies, that there aren't 1000s of other books, that encode the same information.

It's general language modelling capabilities might get an iota better, of course. But for putting factual information into it, wouldn't RAG be a much more solid approach?

It’s unclear whether ISBNdb will scan books without ISBN’s, which were invented in the late 1960’s. Customers appear to be ordering books to be scanned by ISBN? Here is one book seller’s experience:

> Bulk purchases also usually reflect interest in a specific topic, whereas the recent, very large purchases were of books that had little in common, except for the fact that they all had ISBNs. This seller also sells rare books that do not have ISBNs, and none of those were part of the bulk purchases.

Article is paywalled, but I saved a few quotes here:

https://skybrian-links.exe.xyz/post/1026

Is this true? Rare books would very often be out of copyright for a start. What the the actual ruling that says you can scan if you destroy the original? ISBNdb is a database of book information, as you would guess from the name.

There needs to be wider reporting of this if it is true. Where are the journalists ? If this is true, it is scorched earth on the world's knowledge base. And the business models are not proven yet. I can now see why some ai companies paints a dire picture of the future - they are basically annihilating all knowledge sources at the altar of their ai gods. I can also see why there is growing ai-skepticism.

  • Because there's nothing to report. Books like this are routinely thrown out. By publishers, libraries, and regular people alike.

The comments there are absolutely unhinged. There are some good reasons for being anti-AI, but why dilute it with this kind of bullshit:

> It is equivalent to book burning in the past. A form of thought control

  • You're right, it is arguably worse than book burning, not only are they seeking to deprive others of the books, they also want to profit off it.

    • >they seeking to deprive others of the books

      Yeah, this isn't a thing. Throwing out books, even "rare" ones, is a regular procedure. And if you really want to read 'em, you can buy any of these books right now.

  • > why dilute it with this kind of bullshit:

    > > It is equivalent to book burning in the past. A form of thought control

    That's only bullshit if you trust AI companies to serve the book contents without alteration.

From 1984: "Every record has been destroyed or falsified, every book rewritten, every picture has been repainted, every statue and street building has been renamed, every date has been altered. And the process is continuing day by day and minute by minute. History has stopped. Nothing exists except an endless present in which the Party is always right."

I will make a robot scanner for books. I will then scan all the books in my state libraries and make a digital copy of them (without destroying them) before these things come for them.

I wish...

I have zero proof for this, but just a what if: what if Anthropic's strict anti-China stance actually means the Chinese training corpus is way more valuable than people realize?

Maybe we should not have forced AI companies to shred rare books so they can use them for training?

It's not like these books were available to everyone before. If they are destroying one physical copy that's not accessible to the public and replacing it with a scanned copy that isn't accessible to the public, that's not a huge change. Except that will probably last longer digitized. Ideally they wouldn't destroy the physical copies, but this isn't anywhere as bad as book burning in Nazi Germany.

I find going after shadow libraries to be much worse because law enforcement is trying to prevent discrimination of knowledge to the public.

The real blame here should be going onto copyright laws.

Anthropic could take more care by figuring out if the books are still affected by copyright.

But this is just a company trying its best in an unfortunate regulatory environment.

Support your local shadow library:

https://annas-archive.pk/donate

People keep bringing up these rare books but never share what they actually are. What are their names? When were they written? How many copies were in existence? Were they in libraries, or locked up in vaults? Did people have access to them prior to being scanned for AI? How many such books have actually been destroyed?

Weird to see so many of these "trust me bro" twitter stories make it to the front page and cause outrage when no one has any real information.

I can't help but make a comparison to one of those road planning memes.

Just one more book, bro, and we'll solve AGI forever. Trust us, bro, just one more book. Come on, let me have those words and we'll solve AGI forever.

If you consider that AI companies should pay the same for a book, as a person does, you should read about royalties.

Where are the sources for that?

  • The irony in asking for a source here. I guess the idea is that eventually you won't get to ask for a source because they'll all have been destroyed.

    And the only real source for anything will be an LLM response.

I feel outraged, but I also worry this article isn't necessarily responding to what's actually happening.

It's a weird practice to destroy a book when you digitize it, but it's at least an understandable legal strategy to ensure that the digital copy "replaces" the physical one. However, this is only going to apply to books that have active copyrights.

This author suggests that a rare 18th-century botanical text could fall victim to the same fate, but I am somehow doubtful that this is the case. Non-destructive scanning is trivial, and these kinds of books are likely being processed in a quantity that would allow for it without backing up the pipeline.

I really like 404 media, but it doesn't really seem like the evidence points to the conclusion here. Yes, AI companies are shredding books that they digitize, and yes, AI companies are digitizing old, rare books. But the rationale for the book shredding doesn't exist for the old, rare books, so I would need more evidence than just "putting two-and-two together".

At the end of the day, old books with no resale value, including rare books, end up destroyed with some regularity by libraries and bookstores. While this may be an excessively generous take, at least this way the books are getting digitized before they become pulp. The real tragedy will be if the old, rare books that were digitized are never shared with the rest of us, because they were ONLY digitized to train AI, and not to actually preserve anything.

I don't buy this at all. Looking up some of the mentioned rare books, in other sources, on worldcat shows every single book available somewhere. Most of them in dozen or even hundreds of libraries.

So the premise here is that the books are valueless enough that they're being sold by weight, but they're also rare enough that someone will one day wonder where they went, and also that the AI companies aren't also saving the text somewhere much safer than physical media in a warehouse. k.

Why would you need to shred a book from the 1800s when it is in the public domain? Something is fishy here.

Everyone who refused to buy these books in the past was voting for the books' destruction by default. Authors don't owe you a copy of their work. Bookstores don't you real estate to hold a book you don't want and will never want.

The design of copyright has always been to restrict the dissemination of knowledge so that somebody can turn a profit. The handwaved justification is that the profit encourages the creation of new knowledge whose dissemination can be restricted, but that still doesn't erase the fundamental dynamic.

This is merely the latest incarnation. We can imagine a slightly different process on a few fronts - AI companies pay to digitize books (still for their own purposes), but are prevented from destroying the physical copies and they have to openly shared the digitized results. We would view that situation much more favorably - perhaps even as ideal, right?

Those two dynamics could be backed up by court decisions or laws iff they weren't so plainly at odds with how copyright has been and is generally implemented and interpreted. For example, imagine them having to do this through some nonprofit library whose goals was preservation and dissemination. Instead, libraries have been sidelined as things that operate at the edge of the law rather than vital public institutions, whereas shredding books in secret is fully legally condoned.

  • This is the correct take.

    Copyright has done more than anything else to prevent preservation and dissemination of knowledge. And it's forcing Anthropic's hand now. Although they could take more effort to preserve the books.

    • I'm not going to absolve Anthropic here (even though I am a satisfied customer). They could probably get a court decision that it's fair use to non-destructively scan books and keep the physical books, legal inability to directly distribute the results notwithstanding. And they certainly have enough money to pay for some old salt mines or whatever, or fund a library-type institution to do so.

      What I'm indicting is the copyright regime being primarily focused on control and the prevention of dissemination. We can imagine a different world in which the publishers' suit against the Internet Archive went the other way (or was not even brought), and a public interest group sues Anthropic (et al) for destroying cultural commons, and gets a judgement saying all scanning must be done non-destructively and made available through institutions like the Internet Archive.

you wouldn't believe how much shredding your local library does in the name of space conservation and to address changing borrower preferences and demographics. I doubt any AI company comes close to the annual combined library turnover

  • Where I live there's a large library-associated biannual book sale where books are (effectively) reverse auctioned over a period of a few weeks. At the end, anything (with a few exceptions, like the Collector's Corner) that isn't sold is disposed of, with a large 18-wheel truck sized dumpster filled with items to be sent for pulping and recycling. The price at the end is $1 for a grocery bag full of books, so things that don't sell truly are perceived as worthless.

    https://booksale.org/

    Books are information delivery vehicles. We mostly shouldn't care about them any more than we care about a particular set of bits on a disk.

    Publishers also pulp large numbers of books themselves. This is a consequence of the Supreme Court's Thor Power Tools ruling, which clarified tax rules in the US so that keeping large inventories of unsold books was less economical.

> A federal judge ruled the practice is fair use because eliminating the original means only one copy exists at a time.

Really? That sure wasn't a thing when one startup got sued for streaming from its wall od dvd-players, and they adhered to 1 disc = maximum 1 stream at same time.

People would have at least somewhat less of a problem with this if they also put up an archive of PDFs of all these rare books if they are out of copyright.

But that would help competitors with training data, which I assume is why they don’t do this.

Guys never heard of palimpsests and how your regular scans won’t necessarily catch all data there is in the material?

Of course even disregarding this fact, this is utterly disgusting attack on humanity heritage, just as much as any group out there destroying what we should all cherish be it for the historical artifact they represent. Whatever how US judge name it, they don’t worth more than their same-behavior consorts that is terrorists and totalitarian governments.

https://en.wikipedia.org/wiki/Palimpsest

https://organiser.org/2025/07/24/304285/world/china-wages-wa...

https://www.historyexpose.com/things/demolition-afghanistans...

https://link.springer.com/chapter/10.1007/978-3-031-96432-9_...

All those AI powers, and they cannot find a way to do it without destroying the books? Pathetic.

That's exactly why I have been hoarding physical media for years now. I knew they would be buying them up, besides the usual dumpster fire they would end up in when people throw them away. CDs, DVDs, BluRays, books, vinyl, everything. And as they are physical they are not prone to retro-editing. Just like in 1984.

Luckily I don't live in NYC where book hoarders are being evicted because of "fire hazard". Just like in Fahrenheit 451.

This is not that different from having a book burning march. The fervent neocon tech overlords and their funding circles want a monopoly; not only on truth, but on ideas and thought in general.

I don't see any proof of shredding here. Most book scanners I'm aware of are from Google's scanning days, and those had cameras plus page turning.

If we use 'shredding' to mean A book is laid flat, its cover is removed, and then a paper cutter cuts through the binding to create a flat stack of sheets, which are then fed to a sheet feeder, then I could maybe imagine this is better than a page-turning scanner. But, sheet feeding old paper sucks shit, bro. It's not fun.

Upshot, I think we'd like to hear from an anonymous frontier lab employee here to see what's going on -- there are a lot of books in Anna's archive available at considerably less difficulty.

  • The idea here is that the frontier labs are no longer willing to risk using the likes of Anna’s Archive, because they have already been held legally liable for that in the past.

    • Are you referring to the META lawsuit? I think the current landscape is not bad for the frontier labs -- it's settled law that it's legal to 'read' and ingest this data. Anthropic went ahead and just settled a licensing deal for content. The open issue in that META suit is whether or not any distribution happened, as I understand it. I'm certain they all have full backups of the archive somewhere in the org.

      1 reply →

This kind of reads like a blood libel. My guess is they're buying all those books that University libraries are throwing away these days (to see other HM threads for that), kinda sad but probably aren't "rare books" in the way people are thinking

And this is the beginning of the end for content creators.

Why would I spend hours creating original content if Google can extract it and present the answer directly in an AI Overview? What is the incentive to keep doing the work?

If creators stop producing high-quality original material, the information we get over the next few years will increasingly be based on recycled, low-quality garbage.

  • Take a look at most successful journalism today. It's behind a paywall. You get paid by the people who are interested and value your work.

    Can Google steal it and present it in an AI overview? Well kinda. Today Google is doing a trick - they're saying "You can refuse to consent to being fed into the slop machine, but if you do we won't crawl you for Google so you'll get no search traffic. But you're not going to get search traffic anyway! So you might as well opt out of being fed into the slop machine. And companies are starting to do that [1]

    It's really interesting, because essentially what it means is Google is turning into a walled garden, but there's nothing growing inside it so they have to continually import new plants to live in their walled garden and they're going to have to pay to do that. So soon Google will be paying news sites for the right to plumb their feed into the slop machine.

    [1]: https://www.wsj.com/business/media/google-search-publishers-...

  • It’s even noted they’re looking for text pre 2022 as afterward, it’s tainted by their own shit, they don’t believe in the crap they’re making.

In a society that values science and knowledge, preservation of knowledge - and no, slurping text up into an LLM is not preservation - is more important than profits.

These companies are regressive book burners.

  • I love it. Literally burning down civilization to generate losses from slop. This reminds me of the Simpsons Halloween episode where Homer gets a taste for his own flesh and eats himself to death.

    Maybe the Library of Alexandria didn't burn, it was digested.

    The market has solved the what do with excess knowledge problem.

If true this is basically unforgivable. The wanton destruction of history is the stuff of the Third Reich and the Taliban. You simply cannot profess to care about knowledge, culture, or civilization while destroying the physical manifestations of the same.

  • Yes, actually you can. The fetish of physical book worship is a fossil of an age when information storage and retrieval was much harder.

    • I strongly disagree with this take, a book might remain readable thousands of years from now but very few if not zero of our digital data formats likely will. We shouldn't be so quick to throw away diversity in the way information which may be useful for future generations is stored for the long term.

      3 replies →

This is potentially very bad.

Like, Library of Alexandria or Council of Nicaea bad.

We may never be able to recover the information if, say, one of these AI companies copied or translated it wrong then destroyed the source material.

Maybe it's from bad OCR, or maybe from a bad actor - but there are a lot of ways history and information could change in this game-of-telephone like transfer of knowledge.

What is the point of destroying the source material? I don't buy the copyright thing.

  • > What is the point of destroying the source material? I don't buy the copyright thing.

    It is the copyright thing.

    Despite what people say about scanning, the fact is, non-destructive scanning machines have been built and perfected long time ago. This was preferred in the past, back before some major kerfuffle with the publishers during COVID, but that incidentally happened to be before LLMs became a thing, so AI companies never had that option available.

    • You would believe anything hahaha.

      We'd be batteries if you were the spokesman of The People xD

      Even a 10 year old knows they are simply changing information and destroying the source, so that the lie is now in the LLM and you can't prove otherwise.

      They're making the LLM "the source" and destroying the original.

      Only a total midbrain would defend it and be unable to see the obvious.

  • They're destroyed for 3 reasons:

    It is the cheapest way to get them scanned.

    It is the fastest way to get them scanned.

    It doesn't need to be safely archived for another century until it is resold to someone that has not yet been born.

  • > What is the point of destroying the source material? I don't buy the copyright thing.

    The main reason is the machines they use to scan books at scale destroy the books in the process.

I think people are, on the whole, too precious about old things. In the case of books produced after major commercial printing began, I don't believe it is the paper that imbues the book with historical value.

Indeed, I think there's a high chance that this process increases preservation of the most relevant part of the media - the actual content!

There are tons of old rolls of film slowly rotting away in warehouses that were never digitised. Even for beloved media, the BBC occasionally tracks down an old lost episode of Dr Who.

For now these books are in corpuses of training data, but eventually I trust they will make their way to the rest of us.

  • > For now these books are in corpuses of training data, but eventually I trust they will make their way to the rest of us.

    What makes you think they will? What would be the incentives for these companies to do so?

    • Well, you probably weren't going to go and find any rare, non-digitised books to physically go and read (unless you were going to, in which case, rock on), so we can start by benchmarking relative probability there.

      1. At some level of critical information-withholding mass, a leak or disclosure similar to SciHub is inevitable because of the commonly held opposition to hiding knowledge.

      2. Availability via Google Books or similar.

      3. Availability via AI model reference.

      4. Failing any of the above, better AI models that are more capable of doing more things, at the expense of books that were likely to go unread (revealed preference, rare books are often rare for a reason). This will obviously be a nonstarter if you don't want this to happen, but I think it would be good for the world if it did.

      I think category of old books that were going to be read or otherwise become important parts of human knowledge that have not yet been digitised and now will never become so because they are instead being shredded and will never make their way into the light because of AI company data hoarding is a small category.

  • So if the Mona Lisa is part of a model we can burn it?

    • I pre-empted this - my argument does not apply to texts where the physical object is a major part of the historical value of the thing. No, I'm not saying to destroy one of the four remaining Magna Cartas that were meticulously copied by hand. But even if I was, we're only dealing with texts here that are irrelevant enough to have never been digitised already - we tend to digitise most things of value and so the Mona Lisa and Magna Carta would never have been part of this discussion in the first place.

      I am however OK with destroying one of the remaining 50 children's books of which only 300 copies were ever printed in a small town in Ohio in the 1970s as a test run for a failed book which was subsequently never commercialised.

      2 replies →