Comment by ben_w

1 month ago

On the one hand, yes; on the other hand, so much of the training data comes from scraping the web that it feels wrong for them to do what they deny others the right to do.

On the third hand, the settlement Anthropic famously had to pay was for copyright infringement because they didn't actually have the right to even access some of the training data they used, so I can see how this might be compatible with the law.

On the fourth hand, I'm saying that as someone who absolutely isn't a lawyer and sometimes gets surprised when reading about copyright cases that sure sound like they ought to have been trademark cases given my limited understanding.

Companies that didn't give away all their content for free to anyone have actually denied AI companies from training on all their data without paying a fee. Reddit, Associated Press, etc.

For those who chose to give it all away, the ship has sailed, but they did choose to give it away for free to anyone so they can't complain that they succeeded.

  • How exactly other websites “gave it all away”? Also examples you list are websites putting some explicit rule eg in their robots.txt or filtering web crawlers. This is all a reaction to existing situation, so Reddit for sure has been scrapped before Reddit realised what was happening.

  • At what point did the authors whose books showed up in the ai companies training data sets “give it all away” as you claim?

    • If the AI company bought their book, then they didn't give it all away. If the AI company obtained it indirectly like a library or 2nd hand, then the author has already been paid when he first sold it. In either case, he could have refused to be so liberal in sharing it if he didn't want it to be used like that, but he preferred to make some money instead.

      2 replies →