Comment by bashtoni
8 hours ago
Are the Artificial Analysis benchmarks really worthwhile any more?
They really don't seem to match my real-world experience, and based on the comments I see I don't think that match most other people's either.
For example, Opus 5 was at the top for some time. My experience is that it's not noticeably better than Opus 4.8, and it definitely seems worse than Fable 5, which AA benchmarks put behind Opus 5. GPT 5.6-sol and Opus 5 seem pretty interchangeable, although Sol is noticeably better at finding problems in code, particularly edge cases.
I have no faith in these benchmarks. Muse 1.3 shouldn’t even be in the same conversation yet it scores above GPT 5.6 & 6.0