Comment by pu_pe

6 hours ago

From what I understand, the rare books in question are not some historically relevant medieval manuscripts, but rather some relatively recent books (still under copyright) for which there are few print copies available for purchase.

I am not sure that physically destructing one copy of this type of book to preserve its contents digitally is so bad. Pretty much anything that is still under copyright should be valuable only for its content, not for the physical medium it's printed on.

This works only if the digital copy is made publically available. I actually agree with destructively scanning a copy of a work that is in single digits as long as that excellent digital scan is made available.

> historically relevant medieval manuscripts,

These are increasingly available online, btw. Historical research is accelerated when historians have direct access to scans of relevant source material. Not destructively scanned, of course.

> Pretty much anything that is still under copyright should be valuable only for its content, not for the physical medium it's printed on.

Well, there are special collectors editions with the signature of the author and gold pages or whatnot. But the AI companies are probably not using those.

Is slightly modifying some weights in a markov chain somewhere really preserving it's contents?

  • It's not really Markov chain as you need full P(x_t|x_{t-1}, x_{t-2}... x_1) instead of just P(x_t|x_{t-1}).

  • They used to put them all up on books.google.com until they were ordered not to if they were less that 100 years old. Most of the stuff that was accessible was copied to archive.org, and almost all of that stuff ended up on annas-archive and library genesis.

    Consequently, many/most books are more available now than they've ever been. This is mostly a copyright question, not a question of preservation. I say mostly, because those scans, without redundancy and fingerprinting, can be changed and bowdlerized in the future without remaining physical copies as a reference.

    I own about 3500 print books, and started a project to find scans for all of them (that I need to get back to.) Average publication year is probably around 1975, and the bulk ranging from the 1940s to the 2000s. I made it through about 1500, and couldn't find maybe 40, most of them bad. e.g. self-published stuff like "My God Heals, Does Yours?" This was a few years ago, if I went through those 40 now, I bet I'd find half of them.

    I hate what the AI companies are getting away with, but only because they get to violate copyright while being aggressive enforcers of copyright and DRM circumvention laws. Destructive scanning, however, is a cheap way to get a good copy of a book online. If that book were then put in a place where teveryone interested in its contents (or who are just hoarders) can get a hold of it, you'll be able to find a verifiable copy of it 1000 years from now.

    As of now, Russia and annas-archive are just a few points of failure that can erase those words forever. It will be done with armed, uniformed men, and people will claim it is not dystopia, but justice. I don't want the only copy of the text of a book to lie in the interpretation of some privately trained LLM.