← Back to context

Comment by mediaman

1 hour ago

They’re not just trained on human prose. They’re sent to RLHF, and also their language changes as a result of RL on verifiable rewards.

Getting it to write well is really hard because there’s no real way to verify whether it’s good prose or not. You and I can tell, but we can’t write a verifier that codifies our judgment.

Maybe they’ll find a way to improve this, but for now it’s certainly one of the harder problems to solve for LLMs.

Part of it is that I think they also have poor theory of mind, which I imagine is also a hard thing to train it to do.

Why is it hard? Ask it to write professionally in mid-twentieth century style English, and without resorting to the clickbait style of writing.

In any event, other LLMs may not automatically have the problem, and don't even require such a prompt. This is a Claude problem.

  • I don’t think you realize how long ago the middle of the 20th century was.

    • I don't think you realize the prompt actually works. The word "style" does it. I guess you like clickbait too much.