Comment by _djo_
7 days ago
What they could do, if they were truly focused on not taking the most destructive path forward, would be to cooperate on a single scanning company that could pool the resulting text for the others on attractive terms, to ensure that once a book was scanned all the others wouldn't need to scan and destroy the same title. That would substantially limit the damage.
Of course, because each company is happy to burn it all down to beat their competitors and being first to any content is most important, this will never happen.
Somehow people are still failing to understand that transferring the scan from a trust, scanning company, or other external party would constitute a transfer of a copy between distinct legal entities that breaches copyright law in an actionable way.
Doing the scan within their own organization does not because it's within a single legal entity and there is no copyright issue involved.
I'm not failing to understand that, I just don't think it's an insurmountable barrier.
You could overcome it through various means, including helping to set up a clearinghouse for bulk rights, implementing a data clean room approach that models could train on in situ, lobbying Congress for copyright law changes especially around orphan works, using Section 108 of the Copyright Act to set up a specific preservation vehicle like the HathiTrust, and other options.
I find it ridiculous that so many in this thread are acting as though AI companies are simply powerless to do anything but buy up and destroy these rare books.
Make a case that it's more profitable to take any of the courses you suggest and you might be on to something but they aren't. All of those things take time and time is money when you're running a business.
2 replies →