Comment by trollbridge

6 hours ago

We reprint old books after checking out copyrights (for all books, this means pre-1930, but for some (I'd actually say most) it also means ones published up to 1964 and 1973, depending on how the rightsholders did (or didn't) do the renewals).

We use a special guillotine type cutter to cut off the binding and then store the pages in a sealed plastic bag which goes in the archives; they're stored there indefinitely in case the book needs rescanned for some reason. We also keep the original, uncompressed copies of the books on magnetic disks.

We also go out of our way to try to find rare books published in 1931, 1932, etc. so they are ready to go once the copyright expires.

And no, no AI company has ever come to us and asked to run training on all of our scanned copies.

> We reprint old books after checking out copyrights

Who is we?

> then store the pages in a sealed plastic bag which goes in the archives; they're stored there indefinitely

Is that the best thing for archival storage? Like could things like chemical breakdown increase the humidity in the sealed bag or concentrate corrosive chemical vapors? I was under the impression the best environment was an actively climate-controlled environment.

  • The issue is mold. Sealed plastic bags for old books is a recipe for mold. At least they're individually stored, so they would develop mold at independent rates. (And maybe the seal allows airflow.)

    Most books have some mold by the time they're ~100 years old, it's just not enough to cause a serious problem. Sealed enclosures (wrappings, bags, tubs) are a nightmare situation. Even packing books too tighly on shelves accelerates mold growth to problematic levels.

Some source or citation or context is needed here, is this the work of a 2 person no profit or a trillion valued pre IPO company?

> And no, no AI company has ever come to us and asked to run training on all of our scanned copies

The value of most very old books for AI training is very low. You don’t really want your AI training data to start biasing toward outdated writing styles. Most of the valuable knowledge has been covered again in modern texts in more depth and detail.

There is interesting value in old texts and it’s important to have them archived. It’s less valuable for stirring into the giant pot of AI training data, though.

  • If you only care about facts, maybe. Even then I'm sure there are countless facts not described outside of old books.

    I have a hard time believing that text valuable to humans would not be valuable to AI.

    • Agreed, in addition LLMs are trained on a lot of useless data, such as the whole Reddit dataset which has to be at least 90% Reddit garbage, and that doesn’t seem to be a problem. I don’t see why ai labs wouldn’t want to also train on older data

> And no, no AI company has ever come to us and asked to run training on all of our scanned copies

Yet

Why cut off the spines? Isn't that how you end up getting unattributed 'dead sea scrolls'?

  • It's easier to get good quality scans from individual pages than it is from pages in a complete book. Imagine laying a book flat, then the page is distorted in the area around the spine. You can work around this (either by trying to correct for the distortion in software or with clever scanners that position the book more advantageously) but it adds complexity compared to chopping off the spine and just dealing with flat sheets of paper.

  • Easier to send them through a duplex scanner (or put them in a flatbed one if they are fragile). Cheaper than buying the automated ones with the page turning robot arm.

    I have done it at home for my books since the mid-00s.

or just send the books/scans to a country that doesn't recognize US copyright