← Back to context

Comment by ungut

6 hours ago

OpenAI plainly admitted that it is impossible not to do so in a House of Lords inquiry. So, presumably there is no way around it to train models. There is just not enough non-copyrighted data out there.

You mean this one https://committees.parliament.uk/writtenevidence/126981/pdf/ where they write "it would be impossible to train today’s leading AI models without using copyrighted materials"? That doesn't mean they have to download those materials illegally. For a billion dollars, you can easily buy one legal copy of each book in Anna's Archive and still have some cash left over to run a whole-of-internet scraping operation.

  • I wonder if anyone has run the numbers on what the actual cost, both in cash and logistical headache, contacting so many copyright holders would be. That seems like quite the feat to calculate.

    • There's an established network of intermediaries that can supply a large variety of books for a few dollars apiece, so no need to contact copyright holders directly.

      1 reply →

  • I'm pretty sure we would know if they did that. And we don't.

    Plus this is not legal in the EU (and Canada, and ... let's just say the entire rest of the world, and accept that I'll be wrong for one or two smaller countries). Doesn't that matter? Or is only Mistral disallowed from training on copyrighted materials? Je veux ma chaton fat, goddamit!