Comment by chmod775

4 days ago

The more interesting finding is that it's still the second most expensive model (after Fable 5) by a long shot.

At least two models (GPT-5.6, Kimi K3) match its score (~1-2% diff) for half the cost.

The chart shows max effort, used mostly by price-insensitive enterprise users. At medium effort it drops to almost half K3’s cost, and is probably sufficient for 95% of coding tasks.

  • Yeah if there’s one thing that people should really understand it’s that it’s cheaper to have smarter models with less thinking than cheaper models with more thinking.

  • Enterprise users are price sensitive, because they are charged per token (and have limit per user set by companies, and those limits are different from $50 per month to 1500$ per month). Subscription users might be insensitive to that.

  • > At medium effort it drops to almost half K3’s cost, and is probably sufficient for 95% of coding tasks.

    I've also found that Kimi K3 on Max reasoning is benchmaxxing a little bit, High is probably enough for most dev work as long as you have good tooling and a good, detailed plan (which you can create on Max reasoning if you want).

  • For a fair comparison, you should compare to K3 (which AA has not tested yet unfortunately) and GPT 5.6 Sol also on medium or the closest equivalent

    • No, because Opus is smarter than those models if they’re all on medium settings. You should compare at similar levels of performance, which would be favorable to Opus.