Comment by orbital-decay

2 days ago

Most of this article seems like... common sense? Not sure how it's related to the latest generation in particular. I usually find Anthropic's advice on how to prompt their own models deviating from what I see in practice, which is puzzling. Their system prompt was always way too bloated and they kept it as a huge piece for some reason, instead of breaking up into parts. Shouldn't they know better? I wonder if they looked at Pi performing great with minimal amount of distractors in the context and cut their prompt down too, pretending they found something new in their recent models.

> Earlier Claude models could sometimes need repeated instructions or be more likely to listen to instructions at the end of their context window than at the start.

This seems to imply they solved serial position biases like lost-in-the-middle and recency/primacy? Sounds dubious. Labs started claiming this early 2025 and some benchmarks agree, but every time I run an eval on real use cases it's clearly there, especially at longer contexts.

> Most of this article seems like... common sense?

i think you'd be surprised. every model release there's seemingly hordes of people who proclaim the new model is terrible and they're going back to the old one, and it all stems from people still prompting and having their configs setup like we're back in the sonnet 3.5 days

  • I have a coworker that was complaining about Opus 5 and had random shitty skills and custom plugins wired in from YouTube tutorials watched over the past year. He also speaks with the model like it's GPT 4o.

    Needless to say, none of the new models have worked well for him, and he refuses to remove the "tweaks" or update his style of communication, which is obviously breaking the experience.

    • > Needless to say, none of the new models have worked well for him, and he refuses to remove the "tweaks" or update his style of communication, which is obviously breaking the experience.

      All attempts to control the output in a useful way for the user, in a way where the output is as reliable and repeatable as possible... and with a system not at all designed for it, that gets worse the more rules you throw at it.

      Seems like a problem.

      1 reply →