← Back to context

Comment by scosman

8 hours ago

So GLM 5.2/Gemini 3.6 level intelligence for $0.28/m output. And their updated Pro model coming soon....

Plus a size you can genuinely run at home: Unsloth lossless Q8 at 162GB.

I would like to see your "home"

  • Two (linked) DGX Sparks would do it I guess. Though probably slowly (I'd guess 15-20 tok/sec for decode, but higher for prefill). So ~$8-9k USD at current RAM prices, substantially less if they ever (sigh) drop. Electricity use would actually be relatively modest.

    But it makes little to no sense as long as API prices are what they are. Except for maybe privacy reasons.

    • It's at least 2.5x that speed for dual sparks and prefill is good as well. Basically going on vibes it is faster seeming than what one gets by default with openAI or Anthropic.

    • the rational in one’s mind is similar to buying expensive supercar but no driving it daily.

      owning a few GPUs is a lot cheaper than supercars.

      2 replies →