← Back to context

Comment by zamadatix

7 days ago

I'm less concerned with how these rare books didn't rot in large lots of unused books and more concerned with whether or not this helps preserve the content longer.

I seriously doubt any of the companies doing this will preserve the scans for very long. It costs money to store stuff, and none of these companies are have any concern for anyone who isn't them, so they're not going to spend the money or lift a finger unless they work out a way to make it profitable.

  • archive.org estimates it costs $2/GB to store data in perpetuity (well, at least until decades of storage scaling trends stop of course). That's probably less than it costs to acquire and scan the data in, I'd be surprised if they just tossed it out at the end when it's so cheap to store. The same thing happened at the healthcare data lakes I worked on where the trend switched from the usual "how long are we legally required to store this information" to "how long can we legally hold this information".

The content is preserved obviously, and I hope that sometime in the future it will be made available in its original form.

What worries me about the trends is inevitable sanitization of content or straight out falsification.

  • >What worries me about the trends is inevitable sanitization of content or straight out falsification.

    That sounds really speculative, and not inevitable at all.

    I think you're just trying to invent things to be worried about because you don't like AI and don't trust AI companies.

The content is never made available. It would be illegal to.

  • It will be legal to share in some decades. Not that Amazon will bother, though.

  • I don't expect an amazon.com/checkoutrarebooks page but I also don't really think the legality of sharing the content has been much of a concern for these companies either. Ironically, one of the few things Meta got in trouble about with building their AI was helping share the training data.