← Back to context

Comment by benjiro29

10 hours ago

Opus 5 isn't very reliable for coding and introduces a lot of regressions every single time I use it.

And GPT 5.6 Sol over engineers just about everything. No LLM is perfect, its about learning the issues with each LLM and figuring out if you can live with it. Knowledge means that you can anticipate if it tries to pull something funny, and harness it against that behavior.

Sure, but I am a long time Opus user 4.5,4.6,4.7,4.8 and I wonder what's wrong with 5?

  • It seems to have trouble remembering the whole context, even when its limit is only half full. Three times this weekend I've had to switch to Fable, where I literally ask "review the recent conversation and tell me where we went offtrack" and Fable immediately identifies the problems that Opus was having.

    I'm doing data science stuff so it isn't super complicated code; it is about applying valid statistical procedures and techniques. Still, on the code part, Opus 5 had a lot of trouble merging 2 branches yesterday...

    On a tangent, I am beginning to understand why we have replication crisis in academia. I thought C++ was full of footguns; it has nothing on statistics. With statistics, you don't get a compiler error or a crash when you hold it wrong.

  • I remember when 4.7 and 4.8 were released and people were asking what's wrong with them and 4.6 is the best.

    But yes, I also think it's not the greatest model for programming. On the other hand, for agentic tasks that are not programming related it's hard to beat Opus 4.8. It can try different things and pivot even when the user is not great with prompting. 5.0 seems to not be worse, but definitely wastes more tokens and costs more.

this would be _Great_ advice if you owned your own LLM and your knowledge was trapped in Amber because you were satisfied.

It's horrible advice given what we've seen consistent: changing alignments, changing guardrails, changing system prompts, changing inference priorities, etc.

Anyone who relies on these for their work product is chaining themselves to a matrix multiple of indetermintism.