Comment by swiftcoder

2 days ago

> it's the only model in the whole lineup that isn't priced insanely

$4,000 isn't priced insanely? ye gads

Compare to the cost of professional-grade tools in other trades and craft hobbies.

Sure, $4000 can be a lot of if you're a casual hobbyist or are struggle to meet everyday lifestyle costs, but it's definitely not "insane" if this is the trade you make your living from or if you've established a lifestyle that affords disposable income for your hobbies.

And for some people, $4000 for a device you have complete control over and can repurpose and tinker with to your own needs and curiosities is a much much more justifiable expense than a $200/mo rental for some narrow-access tool that somebody else controls.

  • That's only half the reason it's expensive.

    The other reason is that it would likely take years to spend $4000 (plus the real cost of electricity) worth of tokens on a 3rd-party provider that's running a similar limited, DS Flash type model. By that time, the hardware will be obsolete, assuming it's still operational.

    • > it would likely take years to spend $4000 (plus the real cost of electricity)

      Since that cluster only yields 20-30 tok/s on that size of model, at least a decade before the hardware breaks-even with current token costs, and that's not counting electricity. Assuming continued downward pressure on token prices, and the cost of electricity, it never pays for itself.

      18 replies →

    • > The other reason is that it would likely take years to spend $4000 (plus the real cost of electricity) worth of tokens on a 3rd-party provider

      That's just a one-dimensional thought! Your own hardware gives you complete control, and it doesn't time you out for 4 hours, unlike those vendors.

    • But if you can use cloud models, why wouldn’t you use SOTA? For 2400 USD or less per year you can get pretty huge amounts of benefit out of that (though at the whim of whoever you are giving the money to).

  • I am pretty confident that given a $200 subscription on any of the big labs, you're getting $4000-$8000 per month in subsidized tokens... do what you wan't with your dough... and I too have a spark that I got really early (October 2025), but no, economically it does not compare to what's runnable locally in terms of quality from the frontier models. Economically, it looks like for as long as there are subscriber plans, you're better off renting.

    Before getting the spark, I was just using a google colab account, their $49 dollar plan allows you access to h100's and I can run qwen there in a Jupyter notebook... and if I really need that web front end I can just use cloudeflair/tailscale/the local ssh client to reverse tunnel it.

    • This should be obvious but with a model running on local hardware you can do your own RLHF and mod its behavior however you see fit. With cloud hosted models you can't. A few years ago when the models were smaller there were people undoing the guardrails, censorship, and general lobotomization with some form of a RLHF training. You can't do that on larger models unless you have the hardware like this person does.

      Notice all the comments saying like "omg why so expensive so just use the API??". It's a trick for lockin even with, so called, "open" models. Keep trying to run them locally, keep undoing the lobotomies, mod model behavior so that they work for you and do what you want vs only what someone else says they're allowed to do.

      2 replies →

    • Anecdotally, ~$500-1500/month token spend at API OpenAI/Anthropic pricing seems pretty realistic for full-time engineers at companies with "liberal but not unlimited" LLM spend policies.

      This is of course anecdata. I know plenty of outliers, too. I know a principal engineer who uses many multiples of the number I quoted above. I am sure we also know many people making do with much much smaller budgets as well, via all kinds of well-discussed methods.

      But, "$500-$1500 per month per full-time developer" is just kind of the personal mental baseline I use when making my decisions with regards to thinking about whether any of this makes any economic sense.

    • With multiple 200 a month subs you are getting a multiple of those subsidized tokens. At least if you tabulate at retail api prices.

      This rent in the era of expensive hardware thing is not exclusive to inference.

      I’ve needed x86 architecture for windows builds recently and have just hemmed and hawed over buying a decent windows 11 box.

      I can’t make the math work against Azure instances.

      I can spin up a nice one for build deallocate,spin up something cheaper for QA and then turn that off.

      I can build all the devops around that, with a number of passes, with a skills based interface so working with the cloud is not too bad.

      The only thing that still has me thinking about it is the prospect of price is going up even more, which is acid as far as I know.

      And I’m hopefully going to need this x86 stuff enough that I don’t wanna wish I had gotten one for that high prices now.

    • The cloud stuff is definitely a much better economic value, but I would argue:

      1. You learn a lot more running this stuff yourself (especially since you can poke at its internals if you're interested or watch the reasoning chain.) Just being a consumer of this stuff doesn't really teach you much about it other than model & harness specific tricks that become obsolete pretty quickly. (IE, your Claude.md from 6 months ago probably needs a rewrite). Which is fine, I don't think you're going to be "left behind" if you're not a hardcore AI enthusiast or anything (I'm not), but as a guy that's always been interested in computer science I want to see how it ticks.

      2. You can't really depend on this subsidization lasting forever IMO. I know the financials thing has been beaten to death but I guess I'm in the camp that it's good to be in control of your tools so that you can go elsewhere if the economics change.

      I like to check in with ccusage pretty frequently, and honestly like if I were paying API prices for Claude I'd probably be paying thousands a month.

      2 replies →

    • I am also not sure I would choose to use the cheap and easy to run at home model, given a choice. The marketing copy says this is a frontier model, but it's not. Sol and Mythos are the frontier right now. GLM 5.3 Flash simply isn't. I'd rather use the frontier model as they waste less of my time than even Opus.

  • Yeah, in any other profession where you need to buy a van to drive stuff around, you easily spend similar amount of money on capital investment.

  • For $200/mo you either have a SotA model you can’t run on those devices or you have a cheaper model where you pay less than $200 or have a really big amount of tokens without the energy costs and the risk of failing machine

Compared to pricing from 3 years ago, it's insane.

The Sparks admittedly are kind of anemic: 273GB/sec is the same bandwidth as a midrange 4060, although (depending on how you configure things) you can effectively have much greater bandwidth by connecting them.

Compared to 1-2 years worth of LLM tokens for a full-time software engineer making $100K+/year, a one-time spend of $12K for 4 Sparks for on-prem private LLM inference starts looking reasonable, particularly if privacy is an important consideration. It starts looking even more reasonable if running something like a private cloud to service multiple developers because then you likely need less hardware per developer.

(Also, it is going to be a long time until RAM+GPU prices return to what we used to call "normal." If ever. I am not endorsing the current state of affairs and I am not saying you wrong to find it insane, but it is definitely the new reality)

"For new hardware in 2026 with 128Gi of high-speed memory"

Checked a couple days ago and looks like we're at about 3.5x 2020 memory prices (looking at just $/GB).

They only went up from 3000€ to 4000€ which isn't a lot.

For comparison the cheapest Strix Halo 128GB went from 1600€ to 2600€ in the same timeframe.

  • > They only went up from 3000€ to 4000€ which isn't a lot.

    Keep in mind that the DGX Spark was delayed quite a while, meaning that starting 3k price tag is already well into the RAM crisis - just 2 years ago, a 128GB DDR5 kit could be had for $600

> $4,000 isn't priced insanely? ye gads

It depends.

My bicycle was in the 5-digits brand new (now I paid it 1/5th of that and I do thank the first owner for that: the 8 000 out of 10 K I saved were put into stocks, that's his opportunity cost, not mine).

Or I know a great many a going to cry "audiofool", but I can say with certainty the following does sound better than the stereo setup of those crying audiofool:

https://youtu.be/TQg9FTBMcTQ

(not my setup but I've got those speakers: same thing, 15 K EUR brand new for the pair... Previous owner forked the money to buy these brand new and, well, I didn't... And I just hooked them to a wonderful, cheap, fully-integrated Yamaha amp: amazing sound).

If your hobby is DIY job around the house, the cost of tools can very quickly add up too: having 20 K worth of tools is definitely not unthinkable.

You like old cars? Pricey hobby.

Some here even track their cars: tires and brake pads budget (and overall car budget and depreciation)... Through the roof.

There's a saying that you're not really into computers if your setup doesn't cost more than your car.

Is $16 K ($4 K x 4) a lot? It's six months of rent for me and for many here I'm sure. It's not "crazy crazy".

Can anyone afford that? Definitely not. But there are way more insane things out there.

And thanks to the individuals that go through to all the pain of setting those up, we've got feedback, tutorials, explanation, numbers, etc. as to how to run those at home.

For example I helped my brother set up VMs and GPU passthrough and he's now running uncensored models locally and showing me the different answers between the uncensored models and the commercial, censored, ones.

So to GP who bought four of these: we need more people like you on HN, keep it going, blog about it, be "crazy"!