← Back to context

Comment by hailwren

14 hours ago

It has always seemed to me that they're hacking for dopamine response in moderately interested data labelers.

Even when I add multiple prompts into the claude.md file not to be so sycophant sounding and just be blunt, it's responses are full of "the reason it lands...", "that's not X, it's Y" "Your understanding of X — it's better than most people's" or "you already own the right question...".

I don't like that I like it.

  • The most helpful instructions I've found that curb this: "Do not use superlatives. Do not use persuasive writing style."

    I have other more specific ones to avoid talking about things that it's not doing, but those two sentences have covered a lot of ground for me when working w/ Opus models.

  • I have had success in rooting these out by using the correct linguistic terminology for each. Negative parallelisms, tricolons/polycolons, etc. I haven't come up with the proper terminology for all of them.

Yes! The Claudisms do seem to have this slightly uncanny clickbaity feel to them.

  • I always thought it could be because volume-wise, most English prose is probably marketing copy and actual clickbait; so when you train on the entire Internet, you get a troll adept at writing ads. Then people ask AdBot2000 to write a novel and are upset it reads like the next iPhone launch site.

    • Nah, I think this is a common misunderstanding of how LLMs work, where people think that they mimic the pre-training data. Stylistically everything you see is an artifact of post-training, which is from reinforcement learning not from absorbing mass amounts of text. At some point a person or more recently a bot gave a thumbs up to an A/B tested response including em-dashes and claudisms galore.

      7 replies →

    • No, there's no reason chatbot behavior would have anything to do with frequency of text in pretraining.

  • It's more likely that this is from the training data if they're being trained on reams of Internet stuff.

Interesting! My impression was that this was an artifact of RLVR where this slightly preferred writing style got amplified to the nth degree. It's probably some mix.

Given how frequently this kind of punchy-but-vacuous slop gets voted onto the hn front page, the hacking seems to be working.

I assumed they just raw dogged the internet and if you do that, you see way more of that garbage than anything else. It's just that most of us have visually/mentally ignored all of that either via spam filters or just, you know, scrolled passed it.