← Back to context

Comment by drbscl

6 hours ago

Unfortunately, they're full of it https://artificialanalysis.ai/models/claude-opus-5-5#token-u...

It does work out to be a similar cost per task though

You should probably look at the cost/score graph by effort level instead:

https://artificialanalysis.ai/models/claude-opus-5-5#intelli...

It is most of the pareto frontier.

  • Not disputing the increase in quality, just stating that non-cherry-picked benchmarks show it is more verbose at Max effort

    • so don't use it at max? The benchmarks suggest that high/xhigh are more than sufficient to be ahead and a whole magnitude below max with regards to token usage. I'd treat that as an outlier and not how verbose the model is in general (QED I know)

    • Is verboseness the only measure of token efficiency towards overall task completion?