← Back to context

Comment by peri-cl

6 hours ago

Anthropic paying $1.5 billion in fines for downloading Anna's Archive established a moat. They want it to be illegal to pirate books: they can afford the penalties and continue doing it. Just like they want it to be illegal to run local ML inference.

>Anthropic paying $1.5 billion in fines for downloading Anna's Archive established a moat. They want it to be illegal to pirate books: they can afford the penalties and continue doing it.

This seems like a "heads I win, tails you lose" type of argument. If Anthropic was pro-piracy I can imagine everyone getting mad that they're flouting law and want to "steal from artists" or whatever.

>and continue doing it

Source? AFAIK they were caught and stopped. That's why there was the recent story about how they were destroying old books to scan them.

  • [flagged]

    • > This is an unethical as a company may behave, short of killing people.

      This is hysterical. No, format shifting old unwanted books is not unethical. The books still exist, in an internal digital library. If copyright law were to change to allow sharing orphan works some day, Anthropic could share them. But under current law, the books are preserved digitally and used for transformative uses that all Claude users benefit from.

    • did not follow all the details, but my understanding is that some form of copyright law nudges in the direction of destroy after scan?

> they can afford the penalties and continue doing it.

I thought they could've bought just a single copy of each book and use the content to train their models. In that case, it falls into the fair use doctrine and they wouldn't need to pay the fine. And that will be way less expensive than the $1.5B price tag.

  • That's what they're doing now, when they are established.

    But when it was a proof of concept, they were using pirated data.

    Just like Spotify did.

> Just like they want it to be illegal to run local ML inference.

Citation?

  • Over the past decade I've noticed on HN the following order of frequency in choice of words, most common to least:

    1. Citation

    2. Source

    3. Reference

    Long ago in a career based on original research, I/we ONLY used "reference."

    • While it is definitely over a decade at this point (over two in fact), some of this likely comes from the term [citation needed], that originated on Wikipedia, as a cynical backhanded response to unsourced claims. It has become a catch-all. Language and how it evolves is a pretty interesting subject.

      1 reply →

In theory none of them actually got the right to train on illegally downloaded books. Anthropic was simply punished for doing it once.

One wonders if they're still doing it.

  • OpenAI plainly admitted that it is impossible not to do so in a House of Lords inquiry. So, presumably there is no way around it to train models. There is just not enough non-copyrighted data out there.

  • I thought the outcome of that was basically it's legal to train on books, but they acquired the books in the wrong way. If they went out and bought copies of them and trained it would have been fine

  • Of course they are. They have just put on their Swiss Banker suit now and have all sorts of deflection techniques in place such that, of course, "the money has the stamps that says its clean" (when it it really blood money hidden behind a pretty wall).