Comment by TiredOfLife

3 hours ago

Opus 5 is weird. It scores high on benchmarks, but it seems that majority of those who try to use it day to day hate it

If it were a Chinese model everyone would be screaming benchmaxxed.

Seriously something feels really off about Opus 5. I hope they correct it before 4.6 is removed.

One of the things that came out of the decoded reasoning paper was that Claude models had memorized answers to tests but hid this memorization from the user output and pretended to derive the answer properly. It's only possible to cheat so blatantly in closed models where the reasoning is hidden.