Comment by the-grump
8 hours ago
And, regrettably, The Archive lent books regardless of physical possession.
Publishers had accepted the prior arrangement before The Archive decided to push it, if not explicitly then implicitly by not suing.
I'm a believer in The Archive's mission, and I wish they had treated the goodwill they'd accumulated as something worth preserving and not a currency to be spent.
It has been stated by many before me: lending books should have been handled by a separate entity, especially when they removed the physical backing requirement.
Just to be clear, publishers hadn't accepted the "controlled digital lending" (CDL) premise, not even with the one-to-one ratio. Their position was always "first sale ends when the atoms do". There was even controlling precedent: a few years before IA tried their online lending library thing, there was an "MP3 resale" company called ReDigi that had lost on very similar grounds. The publishers suing IA even made sure to sue in the same venue that had decided the ReDigi case so it'd be controlling precedent.
Furthermore, in the discovery for the Internet Archive case, publishers had already found a case where IA had lent out books despite knowing their partner libraries wasn't actually withdrawing loaned-out copies from circulation. The CDL premise was always just a suggestion, and IA would have still lost their case if they hadn't done the National Emergency Library (NEL) stunt or if they'd been sued in another venue that hadn't had the ReDigi case as precedent.
It's important to note that whenever a company decides to sue for copyright, it is often late, because the company is banking infringements up to the 3-year statute of limitations and because building a meritorious case takes time. The lack of a timely lawsuit proves almost nothing about the intent of a publisher with a valid case against you.
The thing is, I don't even think the whole stunt damaged much of the IA's goodwill? I know of a few people who withheld donations to IA, but that was mainly under the assumption that publishers would be getting a billion-dollar damage award that would immediately bankrupt IA and result in it's archives being sold off to Lexis-Nexis or something. The funny thing is, IA wound up settling for a sum so small they had to promise never to reveal it, and the danger is gone, so the only thing people complain about now is just that the NEL stunt maybe pushed them "above the radar" or something.
It's still insane that shredding books for AI training is legal, but this isn't.
The big insanity is tying this to AI. Shredding books is about format shifting; it's a concession hard-won from copyright establishment, which would otherwise be more than happy to deny you the option to convert the media you owned from physical to digital.
AI training happens to be one of the fields exercising that option, but since it's the current favorite topic for people to hate on, here we are.
You don't need to shred books to scan them. They make book scanners that will "rip" a fully-bound book no-problem, and even correct for the curvature of the page and binding to give you an equivalent image. In fact, the Internet Archive specifically built nondestructive book scanners[0] for exactly the purpose of which AI companies are now shredding books. The smart / savvy thing to do would be to buy those machines off IA and use them to read the books they're interested in.
The reason why AI companies don't do this is that they're cheap and desperate for training tokens. Same reason why they have scrapers that will happily overload web interfaces for Git repos following links to everything, even though you can just Git clone the repo with far less stress on the host. The AI people are ultimately there just to pillage as much knowledge as they can as fast as possible. Their scraping practices are slap-dash garbage.
[0] https://ones-and-zeroes.ghost.io/scanning-all-the-books-the-...
1 reply →