Comment by sph
16 hours ago
What protection do LLM search engines have against training off content generated by other LLMs?
Will we get to a point where AI-generated sites make up a majority of the internet, and LLMs are training upon their own regurgitations, with exponential amplification of all their lies and flaws?
Or will the pre-2022 corpus human knowledge be considered the low-background steel standard, and anything after that less and less reliable unless certified that it has been created by a human mind and untainted by hallucinations?
I have a feeling we're already there.
The weird babbling reported from Opus 5 might be a result of either a bad system prompt or bad training data.
> What protection do LLM search engines have against training off content generated by other LLMs?
You're talking about a scenario that won't blow itself up in the next few quarters, so it's of no interest to them.
I've mostly stopped using the Internet to learn new things and have gone back to books from the library. The majority of technical books at the library were published pre-2020s and hopefully, publishing slop physically won't be profitable enough to flood that market, too. Now that the Internet has largely been destroyed by slop manufacturers, whether or not the words are(/were) worth putting on paper becomes a useful discriminator.
Will we get to a point where AI-generated sites make up a majority of the internet
I dunno if they'll be the majority (I suspect we're alredy close to 'yes, and it's already happened'), but I feel very, very confident that they will be the majority, if not the totality, of sites that the vast majority of people see.
They’ll train on prompts and anything else you send in. Many LLM responses are sorta finger printable: I assume this is intentional