Comment by gwd
6 hours ago
Happened to be testing a "review patches on a mailing list" harness I was developing; here are a sample of the latest results, testing 12 patches containing a total of 14 issues:
Opus 5.5: Found 8/14 issues. Total cost: $15.40
Fable 5.1: Found 7/14 issues. Total cost: $66.34
Opus 5: Found 6/14 issues. Total cost: $15.19
Sonnet 5: Found 2/14 issues. Total cost: $19.15
This is a relatively small sample size, but it was both the best and the cheapest.
ETA: NB this is "Equivalent API" cost as reported by claude's CLI; I was using my subscription.
Good data and goes to show that Fable is melting the GPUs and is priced accordingly. I'd guess that cost to serve for Opus 5.5 is meaningfully lower through architecture advances
I just told Opus 5.5 "Perform a code review on the current branch" to see what it would come up with. The results were not inspiring. It told me there were five issues, one of which was a test-coverage gap on line 848 of ProjectTemplateTests.cs. But ProjectTemplateTests.cs is only 160 lines long.
I told it that it had made a mistake in the line number, and to double-check all the line numbers. It responded "You were right to push on this: four of the five line numbers were wrong, and while checking them I found two findings that were overstated."
Then I noticed in the corner of the Claude CLI UI that it was showing "Effort: medium". I'm pretty sure I had set it to high effort before; I don't know when it reverted to medium, but that's another thing that doesn't exactly fill me with confidence.
I'll try again on high effort to see if it does better, but so far I am not impressed with Opus 5.5 on my first day of using it.
Heh I’ve been doing almost the same - back testing against PR comments - and opus 5.5 matched fable 5.1. About $24 for 10 PRs.
I was surprised how much worse Astra did on correctness; I stopped testing with it. Gonna try sol and Luna but low confidence
I doubt your test if you cant even notice that it is not the cheapest with just 4 numbers to compare.