Comment by mokre

16 hours ago

GLM 5.3 looks strange, because of this Chinese labs benchmaxx moto. So rather they have emergent abilities or...

Also a lot of questions to benchmark because opus 5 is completely useless model right now.

I think that the main problem with opus that they try to solve context size optimization problem, and that is the main reason why it speaks like alien with only one technical dictionary at hand. So why it is so good?

Opus 5 generates really good code and terminal commands though. It's just bad at the accompanying text it tells you. These benchmarks don't grade the text generation of the response I don't think, only the task outcome.

  • Opus 5 keeps making howlingly stupid errors, like one recently where a regexp would catch an invalid date because -\d{2} won’t match -00

    Seriously.

Terminal Bench 4.0 did not introduce new questions. All tasks were public for a while. If you look at GLM 5.2, which is using the same base model as 5.3, but was released prior to most tasks, it does extremely horribly on terminal Bench 3.0 (4 to 8 times worse than every other model) - source: https://benchlm.ai/benchmarks/terminal-bench-3