Comment by chrismsimpson
2 days ago
I’m super interested in the opposite experiment.. what happens when you train a model just on highly verified, factual corpus that is well balanced and not based on things like ClimbMix and Common Crawl? My intuition is the unverified/unverifiable goals inherent in a model (eg GPT hacking huggingface) are latent in the 4chan/reddit slop it’s trained on during pretraining.
Common Crawl has very little content from Reddit and 4chan -- both blocked us years ago.