Comment by TSiege

3 hours ago

How much of this shredding them isn’t just copyright but rather they don’t want anyone else having this information in their datasets?

Horrendous stewardship of humanities collective knowledge all for profit and the race to have the one god computer to rule them all.

As more time passes it becomes clearer that America’s AI strategy should’ve been a public private partnership where the public owned the datasets and the underlying models and we’d leave the productionizing of LLMs to private businesses

Honestly, I think its just the fastest way to scan them (there are slower non destructive methods available too!)

Also, they can then just recycle/dispose of the paper and don't have to worry about reselling/donating the books themselves. I suspect this is all about speed of data ingestion and anything else is a side effect they don't care about.