Comment by timpera
13 hours ago
I agree, Opus 5 scoring higher than Fable 5 on Artificial Analysis really makes me question the relevance of these scores.
13 hours ago
I agree, Opus 5 scoring higher than Fable 5 on Artificial Analysis really makes me question the relevance of these scores.
There is a very simple explanation for why weaker models appear to kick sand in Fable's face: Fable cannot be benchmarked because of its batshit out-of-control refusal policy.
If it actually tackled all of the problems it was assigned, it would presumably kick Opus into the weeds.