← Back to context

Comment by ThibWeb

8 hours ago

It was a bit of a silly challenge, wasn’t sure how workable, learned a lot in the process about what actually drives usage / costs, and how to keep both under control

Considering they were your top two models, how did the flash and non-flash versions compare? Did you use them for different tasks?

  • Not OP, but I've used both quite a bit and I think GLM 5.3 (non-flash) is a vastly better model. The flash variant is good for general workhorse agents, but it doesn't seem to reason holistically about code and over-engineers solutions to each specific problem it solves. But if you use GLM 5.3 to write a detailed plan with little to no ambiguity, GLM 5.3 Flash executes it just fine for a fraction of the price.

  • I have a hard time justifying GLM 5.3 these days. It’s slightly better than Flash but rarely enough to justify the much steeper price. We chose to use usage-based billing only so are very sensitive to model price.

  • When text wuality or for pure but adwansed coding is concerned i always pick glm 5.3. The flash is awesome for everything that dosent really matter though.

Worth a comparison with DeepSeek v4.1 flash, if you've got another month to spare!

  • Ill tell you from my personal use glm 5.3 flash was better then deepseek 4.1 flash, also deepseek liked to yap in his reasoning traces soo fucking much, the yapping was fast but the task was so slow to complete...

    • Interesting, I have roughly the opposite experience regarding speed: DeepSeek V4.1 Flash gets stuff done way more quickly for me than GLM 5.3 Flash. (I was getting 300+ tokens/sec with DS vs ~100 with GLM in my testing.) I agree that GLM 5.3 Flash is a slightly better model, but for me at least, it's not a huge difference, and I'd rather have DeepSeek's speed.

  • Yep, I think next month will be on that. It feels slightly better from a few days of use, and in our WIP benchmarking it scores way higher

  • DS v4.1 Flash is roughly equivalent. It's going to get some things right/better that GLM flash doesnt and vice versa.

  • They've got an image of their homegrown benchmark in the post that lists DS4.1.