Comment by Aurornis

7 days ago

The length of this article and the unnecessary visualizations felt like a waste after getting to the end and discovering they won’t reveal anything about the books that were scanned.

The closest they got was admitting that the rare books weren’t anything that someone might care about in the sense that people assume when we hear “rare books”

> As the bookseller who sold them told me, there are not many people in the world who would care about them in the same way people might care about the first edition of Oliver Twist, but that doesn’t mean they’re not valuable.

Okay? But then why exactly where they considered valuable enough to warrant an entire article about them going to a book scanning facility? Without revealing anything about these books I have no idea if they were classic literary works that were underappreciated, or if this was some old guide about How to Use Microsoft Office 97.

There was a more balanced take on Twitter (which I’m unable to find again, because Twitter) from a book seller who said it was more of the latter type: Books that were rare because they were no longer in demand and most everyone had thrown their copies away. Some parts of the media are doing backflips to try to imply that these are cherished literary classics being fed into the shredder to deprive humanity of something valuable, but the book seller seemed happy to be making sales for useless old books that no human was interested in buying.

As if the AI industry is only shredding "crappy" books.

I recently bought a rare book. It was a boat design book written by a famous yacht designer in the 1940s, but it's been out of print for decades and I had to pay $150 for a "fair" copy with missing dust jacket.

I have an interest in older technology and the old ways of doing things (for instance, how do you lubricate the mast of a gaff rigged boat so the gaff jaws don't jam?) I've often found myself reading very old books that have been out of print for a century.

Sometimes I read those books at libraries and I've been the only person to check them out in years (I started doing this back when they still stamped the return date on a card so you could see when it was checked out). Now most of those books have been disposed of by libraries due to yield management software and I've ended up with some of them in my personal collection, but people like me can only save a tiny sliver when most of them are being bought in mass by Sam Altman. When they're gone the knowledge in them is also gone.

We are burning the library of Alexandria and the HN consensus is "those books probably weren't saving anyway."

  • > As if the AI industry is only shredding "crappy" books.

    Do you, or anyone else, have any source suggesting that they’re buying highly valuable rare books and shredding them?

    The kind of $150 rare book that you had to buy from a specialty collector who graded it is in a completely different category. You’re thinking of “rare books” in the historically rare, valuable, and collectible category.

    The book sellers shipping off orders of 1000s of books at a time to these facilities are calling the books “rare” because they may only have 1 copy, not because it’s a collectible with a high price tag.

    • There are a ton of niche books like the one the parent is speaking of which might not be considered hugely valuable by the general public or less specialised second-hand stores, but are both rare and valuable in their niche.

      For instance there were only limited production runs for a set of books on South Africa's participation in the Second World War, and getting hold of them is increasingly difficult. But they have ISBN numbers, so they fall within the group of books being collected and destroyed by AI companies, and as a result might disappear altogether, along with the knowledge inside them.

      8 replies →

  • Please scan this book and put it into Anna's library, or keep it for the future.

    I would be also interested in participating in your costs.

  • Im with you, my take-away from the article was of sadness for all the books that are going extinct because of this.

    Real humans sharing their unique knowledge, packaged in a book.

    • The thing is, I doubt this is even in the top 10 causes of books going extinct. I feel like people often over-romanticise the medium, especially.

  • It would be nice if your local library hadn't sold their copy but I think your beef is with the libraries, or perhaps with the politicians who failed to fund them in that case.

    Actually I expect that a copy of most books probably is still available in a copyright library but it might not be easy to access.

  • Weeding (the library term of art for selecting works for deacquisition) is a very import part of collection management. It's actually regionally and nationally coordinated so the interlibary loan network does not throw away the last copies.

    https://cdlib.org/west/ https://papr.crl.edu https://eastlibraries.org

    • You're fortunate if that's the case for where you live.

      In the UK weeding is at the discretion of the Branch Librarian. Books I checked out 15 years ago have now vanished from the system.

      1 reply →

  • Once you start to realize just many books got thrown out annually before “AI” you have less sympathy. Libraries are one of the biggest contributors to this because there is simply too many books that nobody cares about.

    I suspect folks are over weighting how much knowledge is being destroyed in these books. If someone actually quantified it that would be amazing but as someone who started going to used book sales at a very young age I just have no sympathy. Most books are worthless. I don’t mean that from a text perspective either.

    • Libraries usually offer those books for sale to the general public. I just picked up two from my local library that way.

      AI companies are not doing the same.

      9 replies →

  • Can we copy the books? It seems like in an effort to monetize writing, we've created a system that incentivizes destroying it.

> or if this was some old guide about How to Use Microsoft Office 97.

Materials like that may be very valuable to software / tech / HCI archeologists soon.

I can't get the tools or local know-how to straighten my scythe blade in a country where every cottage had a scythe less than 100 years ago, with the last scythe-native generation rapidly dying out.

Fast forward a civilizational collapse and that M$ Office 97 for Dummies might be as groundbreaking as a Guide to Using Roman Concrete

Anyone who is into books eventually finds out that we're permanently losing them all the time. Like they're thrown out, and lost forever. University libraries throwing out huge collections of out-of-print material to make room for new books, study spaces, even cafés. Municipal libraries turning over their collections. Books that never made it to libraries going out of print, tossed in the trash after yard sales.

The AI companies digesting this stuff is a net win for humanity. And I'm not a fanboy! Ideally they'd upload them to Anna's archive too, but even if they keep it private forever, at least these books live on in some way in the model weights. Thats better than a landfill.

  • It would be really great if the companies who are doing this would commit to placing the scanned files into a public trust that would coordinate with organizations like the Gutenberg project to ensure that the scanned materials enter the public domain on schedule. Publishing encrypted archives with the keys in escrow would be a good first step.

    IMO that would go a long way to resolve any concerns about losing books. I still don't like the idea of extremely hard to find or last prints being actually destroyed for this, but it certainly makes it more palatable.

    • I was thinking that maybe this is not to their benefit.

      Just imagine 20 years from now its hard to get books in print. AI companies can just change the history by altering their model's content.

      I'm not to keen on corporations holding the world's entire print history in AI models.

    • They cannot do those things. The reason they’re scanning the books is because it’s not possible to legally obtain or transfer their digital copies. They have to do their own scans and keep them in house.

      10 replies →

  • I would feel much better about this process if they were uploaded and if it were framed as a knowledge preservation project. This would only slightly increase the cost of the project, but have a huge impact on its perception and its net positive impact.

    Of course, actually benefiting humanity is only a minor, indirect concern for investors.

    • I thought this was the original purpose of Google Books. It's actually a little surprising to me that there's anything left out there worth buying in physical form and scanning that hasn't already been digitized in some way.

  • If or when they go bankrupt or reach agi, they will just delete them. I hope Anna's archive already has them anyways. Apparently openai's newest unreleased model, gpt 6, is capable of continuous training at Inference time, like a person is. That might be enough to delete them

    • And that’s why imo the outrage should be about copyright law not the businesses. I think it’s a hard problem to solve but right now with copyright on new books being the life of the author + 70 years it’s a bit silly.

      2 replies →

  • That is the problem: they are digesting it. They are not creating a new kind of library, where you could say, show me the text of "How to Fix Your Ice Cream Problems". (An actual book I own) It is not being done for our future reference.

    • They are digesting it an incorporating into the weights. You won’t be able to get the exact page but you will (if the model is good) be able to get the knowledge out of it in a likely far more concise, relevant, and certainly more widely accessible than the book sitting on your shelf.

      2 replies →

  • > Ideally they'd upload them to Anna's archive too

    Not only can they not do that, they must scan physical copies because they are forbidden from using digital pirated copies from sources like this.

    Anthropic had a big settlement because they were caught using downloaded digital copies. As a response they’ve ramped up their book scanning and others have followed.

    • Yeah, the very disjointed way that copyright is applied (I would argue, rights that should also transfer across to digital media but generally have not) is contributing to this situation in the first place.

  • This is the right answer but the reason they can’t upload the scans is copyright law as demonstrated by Google having to settle with the publishers and allow them to remove their books and limit free access to 20% of text. The AI companies are essentially compressing the information in a huge swath of books that would otherwise be headed to landfill and making them 1000x more accessible. This is unquestionably one of those instances where capitalism is taking money from rich investors and benefiting the 99%.