← Back to context

Comment by a13n

7 hours ago

Regarding the Pareto frontier and related benchmarks, I have a hard time taking anything seriously that claims that Opus is anywhere near the intelligence of Fable. Are there any benchmarks that haven't just been benchmaxxed that more accurately represent actual usage?

FrontierCode is prob the closest. [1] It's closed source (so no direct benchmaxxing), and it was calibrated by 20+ open source maintainers. It shows Opus 5 (medium), beating out the other reasoning levels by a large margin. i.e. Opus 5 w/ higher reasoning levels actually reduces performance. [2]

However, you'll have to gauge for yourself how closely their tasks resemble your tasks.

1/ https://cognition.com/blog/frontier-code

2/ https://cognition.com/frontiercode