Comment by yipinwong

2 days ago

The analysis is still not compelling for me to switch from gtp5.6-luna to GLM-5.3-flash given

- costs per task $0.05 vs $0.09

- speed 130 vs 88

- where GLM has only 5 more intelligence point: at this point few point is meaningless for most of models

https://artificialanalysis.ai/models/comparisons/glm-5-3-fla...

Been using Luna exclusively since the price drop, and i've been very satified with all tasks from planning, writing code, and other agent tasks. (just change thinking level from low <-> ultra)

---

btw, I did try out Ox Alpha, the coding feels good but still not way better for me to switch to it.

Luna is at a very compelling point on the price/performance curve.

I have found that sometimes a smaller model with max reasoning is actually more expensive than using the next tier model with a lower reasoning effort. It’s certainly faster.

  • Agreed. with "Ultra" (higher than Max), the luna performs really well for my non-metric-backed personal experience

Which has a better monthly plan? Right now Z.ai "Pro" plan (the middle one) is $56/mo if you prepay for a year.

I signed up for their Lite plan when it was only $28 for the whole year (less than $3/mo). Definitely very happy with that purchase!

I’ve been super pro-Luna lately. I really hope that Gemini-Flash-Lite is positioned to compete with it. We all know that Anthropic has abandoned Haiku and it would never be that cheap.

Probably shouldn’t say this here but I’ve been planning to up my $20/mo exploratory ChatGPT subscription to the $100/mo tier as soon as I hit my cap. Between the progress and quality of Luna and their continuous resets, it’s been a few months now that I’ve lived off the $20 tier, frankly waiting for the need to upgrade, credit card in hand.

I’m always trying new models, like many of us here, but the price is just so good for a well balanced, American, hosted model.

  • Funny, I used to use Gemini before Luna as it was "good enough" and cheap.

    For me at this point, most of newer models are capable enough, I focus more on $ and how much I can save.

> The analysis is still not compelling for me to switch from gtp5.6-luna to GLM-5.3-flash given ...

So Luna is competitive because a few weeks ago they did a 80% price drop?

Many here said that 80% drop was not a move against Anthropic but a move against chinese models and your comments indicate that's the case.

I don't understand what would possibly make someone prefer speed over output? You'd rather get wrong bad answers that don't work as well very fast?

In general I really don't mind waiting 5, 10, 40 minutes. There's other things I can look at, other plans or assessments or outputs aplenty stacking up. Its baffling beyond words to me that anyone would take speed over good output. Surely the better output is going to save enormous time in the long run, have better outcomes. What is it that addicts people so much to speed, especially when the difference is between fast and very fast?

  • Background agent services that require no super human vision.

    You might think faster is better when you vibe code or monitoring. But with background agents, the faster the speed, the more jobs it can perform.

    The speed won't matter as much for your personal projects, but if you want to handle enterprise level request, the faster the better.

    Say you queued up all messages for a task in say Kafka, you got workers calling AI agents. You will have thousands of messages to do, and the faster the AI agents can do its work the better you can clear the queue.