Comment by cameronh90
2 days ago
Maybe others have found otherwise, but I find the benchmarks drastically different to real world "feel" of a model, even within the same harness. I'm not sure if this just reflects personal interaction styles, or if it is indicative of benchmaxxing or unrealistic automated benchmarking methodology.
Opus 5 consistently comes at or near the top, but outputs constant unreadable jibberish. Meanwhile GPT 5.6 Luna medium tends to be rated pretty poor on agentic tasks compared to the Chinese lab open models, but I find the latter much more likely to lose track of their own behaviour during a long-horizon task or get stuck in a doom loop.
(This isn't a comment on GLM-5.3 Flash as I've not used it!)
Opus is a yap god. I've found it much, much better with Claude Code's output style set to `Concise` and this:
https://news.ycombinator.com/item?id=49413456
We shouldn't have to resort to this, but it can be mitigated enough that it stays as my daily worker agent. Although, I mostly use Fable to farm out to Opus agents so I don't have as much exposure to what kind of blathering is going on in there.
It's crazy that sometimes I ask Opus 5 to explain what it just wrote to me, and it declares "that was word salad" (its words, not mine, without any hint from me other than "explain it").
[dead]
Yeah, I just ignore the benchmarks at this point. For open-weight models the provider's setup impacts performance so you can have different experience's with the same model at the same quantization from different provider's. Just have to use them on real tasks with your actual harness to really know how they will perform and hope the provider doesn't do something to degrade performance (e.g. update the middleware to a new version with a defect that impairs performance).
For me it's been this: Opus (at least in my experience) is unbeatable at "planning the work". That includes a lot of things, including getting arch. sorted, a chassis/skeleton done. Filling that up and doing the actual "coding," though, I've noticed no real difference between the Claude model and GLM. So it'll be interesting to see how this flash model compares cost-wise to what I'm currently using, which is GLM 5.3, for coding. Looks like it will be reduced even further and might be great if it's faster (and better?) than 5.3 in my real world/personal experience.
(I'm someone who doesn't really care about delays of a few seconds, or even more than few seconds. But if I am trying to notice then sure Claude is definitely faster as well).
I've been using 5.3 since they initially announced it and my gut feel is that it's not as good as 5.2 for agentic tasks. I'm still using it - I don't think it's bad. I'm just not convinced it's better.
idk i think that i spend significant tokens with both to be able to tell 5.3 is way better overall.
https://kommodo.ai/i/IFSUUQYT522uZXWvePZ1
it still fells stupid sometimes and it is benchmaxxed for sure. but its good enough that im building all the hobby projects with it.
I have had the opposite where I felt 5.3 as a stronger model than 5.2. Its feels way more in tune with my code, and does more nuanced edits. Though I do handhold my models a lot, so might fall outside the agentic term.