← Back to context

Comment by vehemenz

2 days ago

That's only half the reason it's expensive.

The other reason is that it would likely take years to spend $4000 (plus the real cost of electricity) worth of tokens on a 3rd-party provider that's running a similar limited, DS Flash type model. By that time, the hardware will be obsolete, assuming it's still operational.

> it would likely take years to spend $4000 (plus the real cost of electricity)

Since that cluster only yields 20-30 tok/s on that size of model, at least a decade before the hardware breaks-even with current token costs, and that's not counting electricity. Assuming continued downward pressure on token prices, and the cost of electricity, it never pays for itself.

  • I don't understand how people don't consider this.

    Plus you're spec'd out of near-SOTA level in months.

    The only reasons to actually do this are a) you have a lot of dispensable income and are a hobbyist/tinkerer, b) you have real, legitimate privacy concerns or, relatedly, c) you're doing something you don't want to get flagged

  • As a counterpoint, my homelab/home-LLM hardware has appreciated in value by about 60% since I bought it.

    Of course, it's not real unless I sell, and the value will eventually go down, but so far I have significant paper profits.

    Also, DeepSeek token prices are continuing to _increase_, not decrease.

    • > DeepSeek token prices are continuing to _increase_

      One increase does not a trend make. And the current crop of models are now undercutting deepseek flash...

      6 replies →

  • > it never pays for itself.

    Exactly; its a development box for fiddling with GPU hardware with a large amount of video-addressable memory. It's not an inference box, really, though it's neat that I can at all!

> The other reason is that it would likely take years to spend $4000 (plus the real cost of electricity) worth of tokens on a 3rd-party provider

That's just a one-dimensional thought! Your own hardware gives you complete control, and it doesn't time you out for 4 hours, unlike those vendors.

But if you can use cloud models, why wouldn’t you use SOTA? For 2400 USD or less per year you can get pretty huge amounts of benefit out of that (though at the whim of whoever you are giving the money to).