← Back to context

Comment by cameldrv

15 hours ago

Yes! The Claudisms do seem to have this slightly uncanny clickbaity feel to them.

I always thought it could be because volume-wise, most English prose is probably marketing copy and actual clickbait; so when you train on the entire Internet, you get a troll adept at writing ads. Then people ask AdBot2000 to write a novel and are upset it reads like the next iPhone launch site.

  • Nah, I think this is a common misunderstanding of how LLMs work, where people think that they mimic the pre-training data. Stylistically everything you see is an artifact of post-training, which is from reinforcement learning not from absorbing mass amounts of text. At some point a person or more recently a bot gave a thumbs up to an A/B tested response including em-dashes and claudisms galore.

  • No, there's no reason chatbot behavior would have anything to do with frequency of text in pretraining.

It's more likely that this is from the training data if they're being trained on reams of Internet stuff.