Two (linked) DGX Sparks would do it I guess. Though probably slowly (I'd guess 15-20 tok/sec for decode, but higher for prefill). So ~$8-9k USD at current RAM prices, substantially less if they ever (sigh) drop. Electricity use would actually be relatively modest.
But it makes little to no sense as long as API prices are what they are. Except for maybe privacy reasons.
2 sparks currently run this model at 60 t/s single session, up to just over 100 t/s aggregate with concurrency of 4.
Going local has as opened up a world of use-cases I never would have entertained the idea of on metered/cloud usage. Privacy is a large part of it but, I also no longer think twice about whether to send a prompt or not based on the psychology of it costing money.
Cached input tokens on local inference are free, so I
don’t care about running sessions up to 500k tokens and hundreds of turns (it’s rarely useful, but DSv4 remains surprisingly coherent up there)
It's at least 2.5x that speed for dual sparks and prefill is good as well.
Basically going on vibes it is faster seeming than what one gets by default with openAI or Anthropic.
Two (linked) DGX Sparks would do it I guess. Though probably slowly (I'd guess 15-20 tok/sec for decode, but higher for prefill). So ~$8-9k USD at current RAM prices, substantially less if they ever (sigh) drop. Electricity use would actually be relatively modest.
But it makes little to no sense as long as API prices are what they are. Except for maybe privacy reasons.
2 sparks currently run this model at 60 t/s single session, up to just over 100 t/s aggregate with concurrency of 4.
Going local has as opened up a world of use-cases I never would have entertained the idea of on metered/cloud usage. Privacy is a large part of it but, I also no longer think twice about whether to send a prompt or not based on the psychology of it costing money.
Cached input tokens on local inference are free, so I don’t care about running sessions up to 500k tokens and hundreds of turns (it’s rarely useful, but DSv4 remains surprisingly coherent up there)
It's at least 2.5x that speed for dual sparks and prefill is good as well. Basically going on vibes it is faster seeming than what one gets by default with openAI or Anthropic.
the rational in one’s mind is similar to buying expensive supercar but no driving it daily.
owning a few GPUs is a lot cheaper than supercars.
I dunno. I bought the Spark in January and it has led indirectly to paid work.
I don't use it for local inference so much. I use it to learn.
I also use it as my daily driving Aarch64 development system.
Aside it's also very cool what else can be done with unified GPU memory, once you realize you have it...
1 reply →