← Back to context

Comment by ACCount37

7 hours ago

Scanning books by taking them apart into singular pages and scanning those pages is faster and cheaper. AI training is a numbers game, so they want faster and cheaper.

What happens to the pages after? No one needs them anymore, so they get mulched and recycled.

That would be the dominant scanning method even if copyright wasn't a thing. But then again - if copyright wasn't a thing, there would be much less need to scan any physical media.

The reason why OpenAI can't just go on Amazon, buy a "digital edition" of a 2018 book and use that is that it would violate the license in ten ways, and then the DMCA laws that forbid breaking DRM on top of it.

Think about that last point for a moment. Our “rights to read” are diminished significantly with digital works as compared to printed works. Right of resale. Right to lend.

In the end, digital publishing just isn’t right and will lead to massive gap in our historical records. They require active curation and cannot be preserved simply by resting on a dusty shelf.

  • Every innovation since the microprocessor isn't worth saving in the grand scheme of things.

    When today's algae evolve enough into tomorrow's sentient creatures, they're really only going to need up to the industrial revolution and should probably stop right before that.

> What happens to the pages after? No one needs them anymore, so they get mulched and recycled.

Strictly speaking, no one needs the Sistine Chapel or the Pietà etc. It would be a shame if they were mulched and recycled, though.

  • Same with the magna carta and the American constitution.

    ChatGPT know them, i'd count that as digitalised why keep the originals?

Machines for non-destructively scanning books were developed and perfected long ago. The destructive scanning is neither technological limitation nor an issue of expedience. It's an issue of copyright law and fair use.

  • Why is the shredding a result of the fair use stuff? I actually don't understand

    • They’re not legally allowed to keep a physical and digital copy at one time because pf copyright law, and fair use doesn’t cover it as an exception.

> if copyright wasn't a thing, there would be much less need to scan any physical media.

Because there’d be much less content created in any media to capture in the first place.

  • Empirically, probably not. We had lots and lots of content before copyright, and people seem to produce lots of content even in jurisdictions with weaker copyright.