Comment by sam_lowry_
6 hours ago
Why don't Google, OpenAI, Anthropic, Facebook & Co defend Anna's Archive publicly?
Coming out would be a bold move for them.
6 hours ago
Why don't Google, OpenAI, Anthropic, Facebook & Co defend Anna's Archive publicly?
Coming out would be a bold move for them.
Anthropic paying $1.5 billion in fines for downloading Anna's Archive established a moat. They want it to be illegal to pirate books: they can afford the penalties and continue doing it. Just like they want it to be illegal to run local ML inference.
>Anthropic paying $1.5 billion in fines for downloading Anna's Archive established a moat. They want it to be illegal to pirate books: they can afford the penalties and continue doing it.
This seems like a "heads I win, tails you lose" type of argument. If Anthropic was pro-piracy I can imagine everyone getting mad that they're flouting law and want to "steal from artists" or whatever.
>and continue doing it
Source? AFAIK they were caught and stopped. That's why there was the recent story about how they were destroying old books to scan them.
> From the start, Anthropic “ha[d] many places from which” it could have purchased books, but it preferred to steal them to avoid “legal/practice/business slog,” as cofounder and chief executive officer Dario Amodei put it (see Opp. Exh. 27).
https://cdn.arstechnica.net/wp-content/uploads/2025/06/Bartz...
Sounds pretty pro piracy to me.
[flagged]
2 replies →
> they can afford the penalties and continue doing it.
I thought they could've bought just a single copy of each book and use the content to train their models. In that case, it falls into the fair use doctrine and they wouldn't need to pay the fine. And that will be way less expensive than the $1.5B price tag.
That's what they're doing now, when they are established.
But when it was a proof of concept, they were using pirated data.
Just like Spotify did.
> Just like they want it to be illegal to run local ML inference.
Citation?
https://news.ycombinator.com/item?id=49076057 ("Our position on open-weights models (anthropic.com)", 1812 comments)
Over the past decade I've noticed on HN the following order of frequency in choice of words, most common to least:
1. Citation
2. Source
3. Reference
Long ago in a career based on original research, I/we ONLY used "reference."
2 replies →
In theory none of them actually got the right to train on illegally downloaded books. Anthropic was simply punished for doing it once.
One wonders if they're still doing it.
OpenAI plainly admitted that it is impossible not to do so in a House of Lords inquiry. So, presumably there is no way around it to train models. There is just not enough non-copyrighted data out there.
9 replies →
I thought the outcome of that was basically it's legal to train on books, but they acquired the books in the wrong way. If they went out and bought copies of them and trained it would have been fine
Of course they are. They have just put on their Swiss Banker suit now and have all sorts of deflection techniques in place such that, of course, "the money has the stamps that says its clean" (when it it really blood money hidden behind a pretty wall).
Those are all law-abiding organizations, which AA is not.
INB4: "Here is one time one of those organizations broke the law". Don't go there, absolute lowest level of conversation.
No gain, all liability. Easier to cut them a check for access to training data and say nothing. Unless legal discovery was performed, the outside world would never know, and the payment records would roll off corporate records through a record retention schedule eventually. Could obfuscate it as a contractor consulting fee ("knowledge management subject matter expert") if you wanted to get tricky, depending on the risk appetite of whomever would receive the funds.
(not legal advice!)
There’s zero liability in a company stating publicly that they support Anna’s Archive. Zero. Free speech protections cover much more egregious statements than that.
Your freedom of speech is your opponents' lawyers' wet dream. Your publicized support for a known piracy operation will not look very good in the court when you get sued by copyright holders.
Maybe they are worried about claims of contributory infringement?
I disagree. Anyone with even a hint of standing will sue, and keep suing. As someone who has to work with corporate counsel often, do not say anything you don't have to say. Only say what is absolutely necessary. Free speech protects you from your government. It does not shield you from civil suits, and the US is extremely litigious.
I'm pretty sure[0] they're all using shadow libraries, and saying things in favor of them would increase their liability.
Furthermore, every pirate wants to be an admiral. None of the big tech companies are actually in favor of any amount of copyright reform. They never have been. There is a huge gulf between "personally benefitting from copyright theft" and "actually wants to legalize the theft". Anthropic still believes they deserve to be paid for their models, they just have this delusion in their head that doing a bunch of computation on stolen data is equivalent to actual human creativity.
[0] OpenAI, Anthropic and Facebook have been shown in court to be using shadow libraries, I don't know about Google.