Comment by andai
1 day ago
I find the Time per Task[0] metric more helpful, because models vary enormously in the tokens required to complete a task. On Time per Task, Gemini 3.7 Flash is Matched with GPT-5.6-Sol, as well as on price per task.
GLM-5.3-Flash takes 7x (relative to Gemini and Sol) per task. So, it's cheaper, if you don't value your time! Don't value real-time workflows, don't value iteration speed, etc. So, doesn't seem very suitable for interactive or agentic work to me.
But having an ultra cheap model for async stuff is always very nice. (Still, the last few weeks feel less about tech and more like a contest between who can afford to give the biggest discounts!)
--
I also like DeepSwe[2], although they measure Output Tokens and Agent Steps, which are misleading when one model has a much faster output speed. (e.g. on their metrics Gemini looks slower, because they don't account for that.)
[0] Time per Task - https://artificialanalysis.ai/?models=glm-5-3-flash%2Cgemini...
[1] Output Tokens Per Task - https://artificialanalysis.ai/?models=glm-5-3-flash%2Cgemini...
No comments yet
Contribute on Hacker News ↗