Happened to be testing a "review patches on a mailing list" harness I was developing; here are a sample of the latest results, testing 12 patches containing a total of 14 issues:
Opus 5.5: Found 8/14 issues. Total cost: $15.40
Fable 5.1: Found 7/14 issues. Total cost: $66.34
Opus 5: Found 6/14 issues. Total cost: $15.19
Sonnet 5: Found 2/14 issues. Total cost: $19.15
This is a relatively small sample size, but it was both the best and the cheapest.
ETA: NB this is "Equivalent API" cost as reported by claude's CLI; I was using my subscription.
5.5 is higher for max effort, slightly higher for xhigh and lower for high, medium and low effort.
The biggest proportional difference seems to be at max (5.5 is 38% more) and at high (5.5 is 21% less).
I think most people run at high and xhigh. At xhigh it is close enough to be task dependent and I don't think most people will notice. At high effort I think it looks like it will be an improvement for most people.
5.5 Max should probably be compared to Fable - it performs a lot better than 5 Max.
parent means that they could get more client / a larger part of the market, which would lead to more income (more tokens) despite lower marginal prices
Happened to be testing a "review patches on a mailing list" harness I was developing; here are a sample of the latest results, testing 12 patches containing a total of 14 issues:
Opus 5.5: Found 8/14 issues. Total cost: $15.40
Fable 5.1: Found 7/14 issues. Total cost: $66.34
Opus 5: Found 6/14 issues. Total cost: $15.19
Sonnet 5: Found 2/14 issues. Total cost: $19.15
This is a relatively small sample size, but it was both the best and the cheapest.
ETA: NB this is "Equivalent API" cost as reported by claude's CLI; I was using my subscription.
Heh I’ve been doing almost the same - back testing against PR comments - and opus 5.5 matched fable 5.1. About $24 for 10 PRs.
I was surprised how much worse Astra did on correctness; I stopped testing with it. Gonna try sol and Luna but low confidence
I doubt your test if you cant even notice that it is not the cheapest with just 4 numbers to compare.
They're claiming a drop in token use too, and that it nets to 40% cheaper.
Unfortunately, they're full of it https://artificialanalysis.ai/models/claude-opus-5-5#token-u...
It does work out to be a similar cost per task though
5.5 is higher for max effort, slightly higher for xhigh and lower for high, medium and low effort.
The biggest proportional difference seems to be at max (5.5 is 38% more) and at high (5.5 is 21% less).
I think most people run at high and xhigh. At xhigh it is close enough to be task dependent and I don't think most people will notice. At high effort I think it looks like it will be an improvement for most people.
5.5 Max should probably be compared to Fable - it performs a lot better than 5 Max.
https://artificialanalysis.ai/models/claude-opus-5-5?models=...
You should probably look at the cost/score graph by effort level instead:
https://artificialanalysis.ai/models/claude-opus-5-5#intelli...
It is most of the pareto frontier.
3 replies →
Disagree. Our internal company tests showed a cost per task drop from 0.35usd to 0.16usd . Opus 5low vs opus 5.5 low
2 replies →
I don't think so, I typically use Opus 5 on High, and 5.5 scores lower on token use:
https://artificialanalysis.ai/models/claude-opus-5-5?models=...
I tested it with Claude Code, and I can confirm it's way cheaper, better, faster and less verbose than Opus 5.
parent means that they could get more client / a larger part of the market, which would lead to more income (more tokens) despite lower marginal prices