Comment by juancn
2 hours ago
I dunno.
It's sad in a romantic kinda way, because of the lost artifact, but the information is what makes the book valuable, not really the medium.
The out of copyright books don't really need to be destroyed anyway for them to be fair use for AI training, and arguable, even if you needed to, you only need one copy per title per company at most.
So it's not a gigantic loss.
The gigantic loss is that they are destroying the original copy.
They aren’t verbatim uploading the text 1:1. They are creating vector embeddings and training documents from it, changing whatever they want since it’s in private and protected by NDA, and destroying the original source of information so that nobody knows what was originally recorded.
But is the information really scanned and preserved for direct access? Or is it just trained into an LLM so that we can only get fuzzy answers about the information?