Comment by hsnewman

3 days ago

I'm sure that local LLM will be far cheaper

Wouldn't be so sure, at least not for a GPU-based system. Quick math for a 5090 running at around 500W generating 100 tokens per second (reasonable for Qwen 3.8-27B) is around 2-3 kWh for a million tokens, which is around $0.60 for some mix of off/on peak electricity rates.

This is in the same ballpark for that same model on openrouter (https://openrouter.ai/qwen/qwen3.8-27b), highly dependent on input/output mix. Deepseek is a much more capable model that you can't run locally on normal hardware, and their rates are insanely cheap ($0.04 / $1 per 1M).

And so far I haven't considered the cost of the hardware. I happen to have a gaming PC that can be put to use on inference when not gaming, but given these numbers I don't think I would buy new hardware to do inference at home. Unless my math is wrong, it seems you're way better off paying for Deepseek than running Qwen or some other locally runnable model yourself. Of course if you have specific privacy requirements or prefer something unique about a particular model you can run locally, the equation changes.

Depends on what level of intelligence you're wanting to use. A vanishingly small number of people can or would want to go to the hardware expense of running something like GLM 5.3 Flash, much less something like K3.

And if you want Astra/Fable/Opus frontier level, then there's no option at all.

But if you don't need that, or you don't need speed... That opens up the discussion. I've been impressed even with how Siri's been doing with the Apple Foundation Models in MacOS/iOS 27 given how small they are.

Edit: I can't even fully spec the M5 Ultra Mac Studio you'd need for GLM5.3 Flash since 512GB isn't available yet, but it's already at $9500 for 256GB RAM.

  • > A vanishingly small number of people can or would want to go to the hardware expense of running something like GLM 5.3 Flash, much less something like K3.

    It's probably worth letting the user specify their actual costs in such a tool. I run a Framework Desktop 128GB that I bought before memory prices got crazy; the current retail price is almost double what I actually paid a year ago.

    • Dear person,

        I am sorry if this looks out of the blue, but I was trying to answer to an old thread, and HackerNews seems to lock old threads and doesn't allow for private messages (unless I am too dumb to figure it out, which is always a possibility).
      

      Anyway, I wanted to reply to the very thoughtful comment you left my after I inquired about your responsibilities at a co/op(1) and I wanted to tell you that I am very proud of you and that you are the type of "hacker" that I aspire to be :)

      I hope this doesn't come as too forward and I wish you anything but the best.

      (1) : https://news.ycombinator.com/item?id=48383220

Best comparison that has occurred to me is the cost a loaf of bread's ingredients might be slightly cheaper than a baked loaf, depending on how you source it. At home you get total control and know what's going in to it. Yet bake at home is still a niche, perhaps a hobby. So I say as someone who's spent hundreds of hours tinkering with local inference, go for it for anyone reading. But most people just want ... some slices of bread, you know?