← Back to context

Comment by ronsor

7 days ago

They are literally not allowed to publish those scans.

Half the reason the books get trashed in this process is because the first sale doctrine keeps copyright from strangling all the freedom in this narrow area.

If they are in the public domain, they could publish them. How many books are in the public domain but have never been digitized and shared in a public archive?

And if the books are not in the public domain, then they should not be allowed to train their AI models with the material without some kind of license or agreement with the owner of the copyright.

  • > If they are in the public domain, they could publish them.

    That'd be an easy fix with new law -- if you are an AI company with book data, you have a burden to make openly available (or require your suppliers to) all public domain book scans.

Depends on the book. I'd wager the majority of the rare books are in the public domain.

But they won't share them because they don't want their competitors to have the data.

> They are literally not allowed to publish those scans.

Emphasis mine. Just provide a digital copy, no questions asked. They will distributed to various archives globally. I'll pay for the drives and shipping. I understand and can appreciate the potential liability, and am willing to launder it to preserve the subject collection(s) and dataset(s).