Comment by ricardobeat
4 hours ago
This page has been in 'New model — collecting baseline data. Degradation detection paused.' state for months now. It seems to never say 'degraded'.
If you look at the graphs, the latest benchmarks are showing a pretty significant dip, and they match pretty well with some horrible experiences I've had in recent weeks. You can see token usage steadily going down, matching exactly what the author measured on his own.
> This page has been in 'New model — collecting baseline data. Degradation detection paused.' state for months now. It seems to never say 'degraded'.
Click the part at the end that says "View historical performance". They wait to collect more data about a new model before adding it to the overall charts.
The overall solution rate continues to climb when new models are considered.
> If you look at the graphs, the latest benchmarks are showing a pretty significant dip, and they match pretty well with some horrible experiences I've had in recent weeks. You can see token usage steadily going down, matching exactly what the author measured on his own.
The y-axis is amplified to make differences look larger than they are.
Hover over the dots to see the confidence interval. A 1-2% change means nothing.
The page is currently tracking Opus 5, which was released July 24 and has not seen any updates since.
30-day average 83% [71-91], last result is 79% [66-88], and it dipped to 75% [61-85] a week ago.
The page says
> We always use the latest available Claude Code release and the SOTA model
They've stepped up the benchmark for each new model. Presumably they're gathering Fable data now.
They also link to Anthropic's public blog post about some degradations, their cause, and how they fixed them. The time period sounds like the "recent weeks" you experienced: https://www.anthropic.com/engineering/a-postmortem-of-three-...
2 replies →