Comment by nullbio
9 hours ago
I've been saying it for a while now, but no one gives a fuck. Let me repeat it again.
THE BIG LABS CLEAN ROOM YOUR DATA (CREATE SYNTHETIC DATASETS ON IT), EVEN IF YOU OPT OUT, SO THEY CAN BYPASS COPYRIGHT LAWS AND THEIR OWN LOOSELY WORDED TERMS OF SERVICE.
"TOS: We don't train on your data" -> Correct. They train on the synthetic version of your data.
I guess we're just going to ignore this forever though. Who cares about the gaping hole that exists in copyright and contract law now that never existed before LLMs were a thing.
For sure. Even if it wasn't a measure to avoid copyright, you pre-process LLM training data to remove errors, characters that can't be tokenized, etc etc. Doing so with another LLM has been standard for a while.