Comment by spwa4
3 hours ago
But they have been training on copyrighted data since GPT-2 at least. 2019, and that's when it came out, so before that of course.
3 hours ago
But they have been training on copyrighted data since GPT-2 at least. 2019, and that's when it came out, so before that of course.
GPT-2 was trained using data scraped from the web (https://cdn.openai.com/better-language-models/language_model... section 2.1), i.e. copyrighted data provided free of charge to anyone with an internet connection.