Comment by yorwba
5 hours ago
AI companies legally acquiring books have indeed been in the news: https://news.ycombinator.com/item?id=49330742
And where are you getting the idea that Mistral doesn't train on copyrighted data? There's not a lot of code written by people who've been dead for more than 70 years, but somehow Mistral has been able to release coding models anyway.
But they have been training on copyrighted data since GPT-2 at least. 2019, and that's when it came out, so before that of course.
GPT-2 was trained using data scraped from the web (https://cdn.openai.com/better-language-models/language_model... section 2.1), i.e. copyrighted data provided free of charge to anyone with an internet connection.