Comment by yorwba
10 hours ago
You mean this one https://committees.parliament.uk/writtenevidence/126981/pdf/ where they write "it would be impossible to train today’s leading AI models without using copyrighted materials"? That doesn't mean they have to download those materials illegally. For a billion dollars, you can easily buy one legal copy of each book in Anna's Archive and still have some cash left over to run a whole-of-internet scraping operation.
I wonder if anyone has run the numbers on what the actual cost, both in cash and logistical headache, contacting so many copyright holders would be. That seems like quite the feat to calculate.
There's an established network of intermediaries that can supply a large variety of books for a few dollars apiece, so no need to contact copyright holders directly.
This is very true. As someone with quite the experience with materials published under Penguin, Scholastic, etc. you effectively have a "dictionary attack" on the matter, rather than true "brute force," but that still leaves quite a list to compile to send to each and is easier for larger titles than smaller ones. I wonder how that leads to a bias in what materials get used for training. You are not getting many local self-published books this way.
It is almost like we need a "for use for training" agreement across the board. This would not fix the current issues (at least without substantial work), but going forward would allow for creators (or publishers/rights holders) such as this to designate a work as crawl-able for AI. A robots.txt just for Claude.
Doesn't really matter. The incentive structure to steal clearly exists, so why would they even go through the trouble?
Pretty easy to assertain that they don't acquire them legally due to the plethora of evidence and court cases against them. No copyright holder would be sueing them if they knew they sold the works in the first place.
I always wonder why y'all feel the need for these impressive mental gymnastics. You can use the models /and/ think they are trained unethically. Living through the ambiguity without abandoning your ideals completely is a valuable skill these days.
I'm pretty sure we would know if they did that. And we don't.
Plus this is not legal in the EU (and Canada, and ... let's just say the entire rest of the world, and accept that I'll be wrong for one or two smaller countries). Doesn't that matter? Or is only Mistral disallowed from training on copyrighted materials? Je veux ma chaton fat, goddamit!
AI companies legally acquiring books have indeed been in the news: https://news.ycombinator.com/item?id=49330742
And where are you getting the idea that Mistral doesn't train on copyrighted data? There's not a lot of code written by people who've been dead for more than 70 years, but somehow Mistral has been able to release coding models anyway.
But they have been training on copyrighted data since GPT-2 at least. 2019, and that's when it came out, so before that of course.
1 reply →