← Back to context

Comment by greenlimetea

4 hours ago

This is potentially very bad.

Like, Library of Alexandria or Council of Nicaea bad.

We may never be able to recover the information if, say, one of these AI companies copied or translated it wrong then destroyed the source material.

Maybe it's from bad OCR, or maybe from a bad actor - but there are a lot of ways history and information could change in this game-of-telephone like transfer of knowledge.

What is the point of destroying the source material? I don't buy the copyright thing.

> What is the point of destroying the source material? I don't buy the copyright thing.

It is the copyright thing.

Despite what people say about scanning, the fact is, non-destructive scanning machines have been built and perfected long time ago. This was preferred in the past, back before some major kerfuffle with the publishers during COVID, but that incidentally happened to be before LLMs became a thing, so AI companies never had that option available.

  • You would believe anything hahaha.

    We'd be batteries if you were the spokesman of The People xD

    Even a 10 year old knows they are simply changing information and destroying the source, so that the lie is now in the LLM and you can't prove otherwise.

    They're making the LLM "the source" and destroying the original.

    Only a total midbrain would defend it and be unable to see the obvious.

They're destroyed for 3 reasons:

It is the cheapest way to get them scanned.

It is the fastest way to get them scanned.

It doesn't need to be safely archived for another century until it is resold to someone that has not yet been born.

> What is the point of destroying the source material? I don't buy the copyright thing.

The main reason is the machines they use to scan books at scale destroy the books in the process.