Comment by ehsankia

4 hours ago

> late last year/early this year

That's an eternity when it comes to coding models.

In my personal experience, we've had almost a step change every ~3 months this year, at least for bigger one-shot tasks. For example looking at Gemini Flash 3.0 vs 3.5 vs 3.8, it went 5% -> 30% -> 75% on DeepSWE, all since the start of the year.

In one sense you're right.

In another one, Opus 4.6 level already solved 90% of my work-day tasks, so while better models have been instrumental into handling a higher % that does not mean that defaulting on cheaper models can't be good.

I run DS 4.1 flash daily, and then cross check with gpt-6-astra and I've nuked 90% of my AI monthly bill while having higher limits and better performance/intelligence than I did just at the beginning of this summer.

tbf, I the happiest I've been working with claude is late last year/early this year (before March)...

  • That's because Opus 4.6 was the last good assistant model.

    Everything after it might be more "intelligent" but is super tuned around end-to-end task (and related benchmarks), not to act as an assistant.

    Now it's *you* being the assistant, reviewer, etc.