Comment by respectattentio

2 days ago

it's for sure better than deepseek flash 07/31

That is saying a lot if Ox Alpha is also small and relatively cheap computationally. I hope so; I love deepseek-v4-flash-0731 and use it frequently. Fast inference is good and fits with my dev style: I like to be in the loop, not let an agent code on its own for long periods of time.

  • From their blog post, it's 320B total parameters and 18B active parameters, so a similar size, but slightly bigger.

    Regular pricing is $0.15 input, $0.50 output... but currently 50% off, making it $0.075 input and $0.25 output. That beats most of the V4 Flash providers, but not all, and obviously tokens per task may not be equivalent.

    I've also just noticed the blog post reveals the Artificial Analysis score - it's a 57, so it's Opus 4.8 / 5.6 Terra level.

    https://z.ai/blog/glm-5.3-flash

It is not going to be cheaper though. I may choose the cheaper one in the end because performance will be marginal, both being flash.