Comment by impossiblefork

4 days ago

How would that be justified, because as I see it would necessarily treat model outputs as more protected than actual human output?

Surely Anthropic/OpenAI ToS isn't something which can have more weight than, let's say, the ToS of a website expressed through robots.txt, considering that for the length, actual human effort has gone into writing the text, at a cost much greater than anything output by an LLM. So so long as the AI firms are allowed to crawl and then train on human-generated websites that don't allow crawling, but are crawled anyway, I think this kind of thing is really hard to justify.