← Back to context

Comment by fellowniusmonk

3 hours ago

I have some internal tests I use for areas where one particular solution/paradigm is dominant but worse.

Opus 4.6 is the last model that's actually useful and can "adjust" its perspective to use the newer & better solution.

Where Opus 4.8-5 has over fit training on worse/older but "dominant" solutions it refuses to adjust.

Not only does this create an existential threat to adopting progress but it also means that if you have a code base that has rare but real world tradeoff the newest versions of Opus 4.7, 4.8 and 5 are worse than useless and become a major dev timesink.