Comment by energy123
21 hours ago
None of my use cases require frontier capabilities but I still pay $200/month to a frontier lab. I value the additional time saved at more than $200/month. If I had to pay actual API rates, then I'm not sure what I would do, but it would not be an easy decision.
Sure, but anthropic is charging businesses based on usage now and tried hard to pull Fable from the consumer subscriptions before Sol and K3 dropped.
Even now on the $200 plan I use up my Fable credits in a single day and had to start using codex and openrouter for more usage because Fable burns $100s an hour when billed on usage.
I thought similarly until I decided to try DeepSeek.
It became an easy decision, even the $200/month by Anthropic sounds like a bad deal.
Yes, the reckoning here will happen in a year or two when the (probably subsidized, maybe?) coding plans become either unavailable or much more costly.
It's also already the case that in larger companies people do not have access to these plans as employees and must use API rates.
There's also the political angle. When Anthropic and OpenAI held back their top of the line models because of Bessent & Trump's bullshit, and threatened to deny unwashed foreigners like me access... I dropped my Codex plan and made do purely with GLM 5.2 for three weeks before OpenAI finally released 5.6 Sol. Feels inevitable that this will happen again.
Or, somebody will come up with a way to serve e.g. Kimi K3 or the new Qwen model in an extremely cheap way. Or DeepSeek releases a competitive model at their cut-throat rates. And then the cost argument just wins.
k3 costs will go down at least 3x within a week of the weights dropping.
we'll get new quants, dspark speculators, distills and optimized kernels
as long as there are near frontier models available there will be inference providers selling them at or below cost of inference in attempt to get market share.
I have not seen that kind of significant drop with GLM 5.2 yet? so curious why you think it will happen for K3.
This is a very large model. Much larger (3x) than GLM. The resources to run it are very expensive.
2 replies →