Comment by alexpotato
5 hours ago
I've been using GLM-5.3-Flash on Ollama Cloud's $20/mo plan and using OpenCode's free models (mostly Muse Spark 1.3) to do a LOT of work and I would say that I hit a daily or weekly limit MAYBE once a month.
My usage plus reading about how tokens just keep getting cheaper and cheaper sounds like a great thing for "the rest of us" but not sure how the frontier AI labs are going to pay back all of their debt if this is the case.
(I get the inference is currently very profitable but if it's a race to the bottom on token pricing, even that won't last much longer)
No comments yet
Contribute on Hacker News ↗