← Back to context

Comment by fallingbananna

18 hours ago

Those 27.3% are still in the ballpark of modern models:

- Sonnet 5 - 12.4%

- Luna - 17.3%

- Grok 4.6 - 20.3%

- Sol - 37.3%

- GLM 5.3 - 41.8%

- Opus 5 - 51.8%

GLM 5.3 looks strange, because of this Chinese labs benchmaxx moto. So rather they have emergent abilities or...

Also a lot of questions to benchmark because opus 5 is completely useless model right now.

I think that the main problem with opus that they try to solve context size optimization problem, and that is the main reason why it speaks like alien with only one technical dictionary at hand. So why it is so good?

  • Opus 5 generates really good code and terminal commands though. It's just bad at the accompanying text it tells you. These benchmarks don't grade the text generation of the response I don't think, only the task outcome.

    • Opus 5 keeps making howlingly stupid errors, like one recently where a regexp would catch an invalid date because -\d{2} won’t match -00

      Seriously.

  • Terminal Bench 4.0 did not introduce new questions. All tasks were public for a while. If you look at GLM 5.2, which is using the same base model as 5.3, but was released prior to most tasks, it does extremely horribly on terminal Bench 3.0 (4 to 8 times worse than every other model) - source: https://benchlm.ai/benchmarks/terminal-bench-3

Astra is 58%. The current title says it's "rivaling Astra"

  • It is rivaling Astra, on their own benchmark that they made (FrontierCode), that they ran themselves in their own closed-source ecosystem that isn’t reproducible by anyone.