← Back to context

Comment by torginus

9 hours ago

hasn't IP law passed the statute of limitations? As in most models are probably trained on output of other models, as creating enough data otherwise is not feasible. Additionally, they are trained on github repos made since the AI boom, which were generated by models with IP issues (who knows what and how).

Thus training on 'clean' data is like trying to unscramble an egg.