Comment by est31

6 hours ago

> You can reprint a bestseller. You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal. So it's going to accelerate.

Aren't they shredding only the books still under copyright protection? How is an 18th century botanical text still under copyright?

IDK about the shredding, it's not nice, but it's more a problem with copyright law than AI companies.

Scanning books you own should be legal from a copyright point of view, and not require shredding.

Second, one should think about abandoned property provisions for copyright works published more than 50 years ago and in danger of being forgotten: once challenged, either you as the owner have to prove that the work is preserved for future generations (e.g. in various libraries around the world), or you have to authorize further copies, or you give up copyright on the work.

Scanning books by taking them apart into singular pages and scanning those pages is faster and cheaper. AI training is a numbers game, so they want faster and cheaper.

What happens to the pages after? No one needs them anymore, so they get mulched and recycled.

That would be the dominant scanning method even if copyright wasn't a thing. But then again - if copyright wasn't a thing, there would be much less need to scan any physical media.

The reason why OpenAI can't just go on Amazon, buy a "digital edition" of a 2018 book and use that is that it would violate the license in ten ways, and then the DMCA laws that forbid breaking DRM on top of it.

  • Think about that last point for a moment. Our “rights to read” are diminished significantly with digital works as compared to printed works. Right of resale. Right to lend.

    In the end, digital publishing just isn’t right and will lead to massive gap in our historical records. They require active curation and cannot be preserved simply by resting on a dusty shelf.

    • Every innovation since the microprocessor isn't worth saving in the grand scheme of things.

      When today's algae evolve enough into tomorrow's sentient creatures, they're really only going to need up to the industrial revolution and should probably stop right before that.

      2 replies →

  • > What happens to the pages after? No one needs them anymore, so they get mulched and recycled.

    Strictly speaking, no one needs the Sistine Chapel or the Pietà etc. It would be a shame if they were mulched and recycled, though.

    • Same with the magna carta and the American constitution.

      ChatGPT know them, i'd count that as digitalised why keep the originals?

  • Machines for non-destructively scanning books were developed and perfected long ago. The destructive scanning is neither technological limitation nor an issue of expedience. It's an issue of copyright law and fair use.

  • > if copyright wasn't a thing, there would be much less need to scan any physical media.

    Because there’d be much less content created in any media to capture in the first place.

    • Empirically, probably not. We had lots and lots of content before copyright, and people seem to produce lots of content even in jurisdictions with weaker copyright.

> You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal.

An 18th century book would be out of copyright so why would it be illegal to keep the original and scan it?

> Aren't they shredding only the books still under copyright protection? How is an 18th century botanical text still under copyright?

It's cheaper to scan the books if you do it destructively. Cost. That's why they're shredding irreplaceable texts. Nothing to do with copyright.

https://www.404media.co/ai-companies-are-buying-tons-of-old-...

  • One doesn’t need to pulp the pages after scanning though. After scanning, they could be rebound and put into a library.

    • That would be a clear case of copyright infringement under current law. You can’t make a copy of a book and then give the original to someone else.

      3 replies →

Old rare books where there are single digit copies should enjoy some sort of patrimonial protection just like museum pieces. You can own them but have the state have the option to buy it if you’re about to significantly deface it or destroy it.

  • The only issue I see is, how could you tell which are those books?

    • You can't, without spending a fortune to investigate the scarcity of some of the titles. I do a little work in this space and I don't know of anything destructively scanned that has zero other copies, but definitely some with single digit known copies, where no previous scan exists. Now there is one less physical copy, but a digital copy that none of us can access (except by tricking an LLM).

      It's not just the big players buying up these archives, either. There are a lot of smaller players, especially in the OCR space, who are buying up huge swathes of works in languages which have much smaller digital footprints, e.g. Arabic.

    • We manage to do this for endangered wildlife too without anyone counting every single specimen; why shouldn’t we be able to estimate how rare a book is?

      5 replies →

Books that are shredded can’t be scanned by competitors.

  • Yeah, this is the point. I don't understand the bulk of this conversation. Copyright doesn't matter, the books themselves don't matter. All that matters is that their corpus of training data grows faster than their competitors.