Comment by cregy

4 days ago

They will need to do something new here, somewhere between patents and copyright, to help protect US intellectual property on the AI stage. Otherwise, US r&d will fold on this. Longer term, 2 years plus, it will be a race to the cheapest as almost all models will be better than our current frontier ones.

American AI companies pirated hundreds of thousands of books, scraped the web against the express wish of site owners, extracted all public GitHub repositories without the explicit consent of code authors and distilled all of that information, basically the entire knowledge humanity built over thousands of years, into large language models. The way these models were built was ethically questionable. The models are out there now, and time can't be rolled back, but I still don't think the companies who built the models by ignoring intellectual property rights deserve any special protection of their own intellectual property.

How would that be justified, because as I see it would necessarily treat model outputs as more protected than actual human output?

Surely Anthropic/OpenAI ToS isn't something which can have more weight than, let's say, the ToS of a website expressed through robots.txt, considering that for the length, actual human effort has gone into writing the text, at a cost much greater than anything output by an LLM. So so long as the AI firms are allowed to crawl and then train on human-generated websites that don't allow crawling, but are crawled anyway, I think this kind of thing is really hard to justify.