← Back to context

Comment by giancarlostoro

4 hours ago

Call me crazy but:

VRAM & Memory Requirements by Precision

• FP16 (Full Precision): Requires ~1,664 GB of VRAM (e.g., an 8x B300 288GB cluster).

• INT8 Quantization: Requires ~832 GB of VRAM (e.g., 8x H200 141GB).

• INT4 Quantization: Requires ~416 GB of VRAM (e.g., 8x A100 80GB)

VRAM aint cheap, Sam Altman ruined the cost of memory, Nvidia doesnt make enough consumer GPUs letting the market go insane over them, I still have friends on 1070s or 1070 TIs because GPUs have been severely overpriced for too long. I remember when a gaming PC was only $1000.

Even so why would anyone not sleep on a model they cannot run?

Seriously, if a single politician stepped forward and said "i'll bring down ram prices" they could then shoot a puppy and call me a slur and I'd still go out and doorknock for them.

Memory companies have price fixed multiple times. They've paid hundreds of millions in fines. wikipedia even has a page on it. https://en.wikipedia.org/wiki/DRAM_industry_price_fixing.

Look at the financials of these companies, they're all making obscene margins and do they plan to increase production? No. Micron is doing a stock buy back to pump the price of their share.

The Micron CEO just recently said this is the exact plan https://www.theregister.com/systems/2026/10/01/ram-supply-se...

There's sanctions, tarrifs, and a DOJ who doesn't give a shit. Until we can fix that the insanity will continue. Phones will be unaffordable. Laptops will be obscene. Gaming consoles will be thousands of dollars. Desktops will be dead.

If you're waiting for some David Ricardo equation to happen, tough cookies, it's not coming.

The market is legally locked down and we're in hostage pricing mode.

And what's the story? You can't afford electronics because we're using it to build robots to take your job? I mean ...

Nobody is coming to save us. That's our job.

  • > do they plan to increase production? No.

    Micron has 3 brand new fabs currently under construction, 2 Boise, 1 in New York as the first of 4 planned for a campus.

    Plus expanding other existing facilities.

    These things take ~3-5 years from breaking ground to full production. You'd have had to anticipate the current demand years before it happened in order to be bringing production on-line before 2030 or so.

    Samsung and HK Hynix also have fabs under construction and planned.

    CXMT started 11 years ago and only now is reaching any real volume. If they decided a year ago to react to the current demand cycle they'd be 6-7 years out.

    Not much you can really do to wish for more fabrication to exist on any timeline not measured in fractional decades.

    Could they do more and react quicker? Probably, but everything I've read on the subject seems to point to 3 years is absolute bare minimum if you happen to have a shovel ready project with the land bought, local permitting completed, infrastructure extended to the site, and a skilled workforce already in place. They could suspend buy-backs/dividends today and dump it all into building production and there would be no material impact until around 2030.

    > The Micron CEO just recently said this is the exact plan

    CEO simply stated the demand pressure will not go away through 2027, and supply will not increase until around 2028 when currently under construction fabs start shipping volume. The article does not support your statement.

    • > CXMT started 11 years ago and only now is reaching any real volume. If they decided a year ago to react to the current demand cycle they'd be 6-7 years out.

      It's taken them this long to catch up to the DDR5 standard. They've only recently been through qualifications to be a DDR5 supplier for the big boys.

      > Every Major Motherboard Maker Now Validates CXMT DDR5

      https://www.techtimes.com/articles/321572/20260725/every-maj...

      After their recent IPO, they have more than enough cash to ramp up in a major way.

      It's just a matter of time.

    • The second Micron boise fab hasn't even broken ground yet, they are still working on the first one. So don't expect these things to be completed in parallel.

      Some of my family is pretty happy, though, with the job security as they are pretty convinced these projects are all going to take much longer than what's being stated publicly. Micron is saying the first chip from the new fab will be in 2027... though they also predicted it'd be 2026. The date seems pretty slippy.

      1 reply →

  • >Seriously, if a single politician stepped forward and said "i'll bring down ram prices" they could then shoot a puppy and call me a slur and I'd still go out and doorknock for them.

    How many people, outside of tech geeks and megacorps care about RAM prices? And how gullible would you be to BELIEVE the politician they could actually make it happen, and even if they did, that it would extend to the average person, and not JUST megacorps/megadonors?

  • RAM manufacturers are bidding against NVIDIA and everyone else for the same constrained supply of EUV machines. And it takes years to build more fabs. Micron has multiple fabs coming online in 2027 and 2028.

  • If we take some time to understand how HBM memory is manufactured (with particular focus on yield risk for final packaging steps), we will hopefully learn that the current capacity crisis is not bullshit.

    I guarantee Micron & friends are not intentionally orchestrating their business such that they would suffer a massively reduced chance of yielding on a per-die basis. Unless someone is actually buying HBM devices, they are not going to be making them. These are not a commodity that can be speculatively manufactured in any economically rational way.

  • > they could then shoot a puppy and call me a slur

    I know it's just a figure of speech, but damn. I laughed out aloud in public just reading this.

  • [flagged]

    • I don't think RAM vendors have formed a cartel and I think this is knee-jerk anger without any thought. RAM is a commodity product with massive upfront capex costs, and those always have boom-and-bust cycles. At various points in the 2010s and 2020s RAM vendors were getting eaten alive by a supply glut, this would not have happened if they were a cartel.

      Is it really so hard to believe that RAM prices are up because demand is simply exceeding supply, especially in a market where additional supply takes years and billions of dollars to come online? There's no need to posit cartel behavior and a fair amount of evidence that there is none.

      5 replies →

There's no BF16, original full quality weights are quantized already and 510GB.

Then good portion of those weights are n-grams (~200GB) that don't need to be in VRAM.

Then KV cache of that model is super lightweight at ~1GB per 1M tokens. If HBF succeeds, then accelerator with 16GB of VRAM and 1TB HBF/NAND is probably all you need (?).

Are you counting the n-gram/PLE as part of the model weights there? They can go in host memory. Would be good to show your working. Also the released weights are pre-quantised and presumably QATed, so your "Full Precision" and INT8 are simply not a version of the model that actually exists.

Edit: I went and checked for you. The LM backbone is 307.2 GB (286.1 GiB), straight from DeepSeek's upload. The n-gram table is 203.1 GB (189.1 GiB), which goes in host RAM. Note the embeddings are higher precision than the expert tensors, so it's a larger fraction of the bytes than it is of the parameters.

So,

> Call me crazy but:

You're crazy. :-)

Projects like DwarfStar https://github.com/antirez/ds4 really lower the hardware bar a lot so Deepseek 4.1 flash and other mixture of expert models can run on consumer hardware. There are also other inference providers who make their money serving openweight models. Services like OpenRouter make it all too easy to utilize these models. Access to these models isn't hard. The hardware moat is becoming pretty easy to bridge.

  • More concretely DwarfStar M5 128GB Deepseek 4.1 flash 1K tokens @ 29s, 5K tokens + reasoning @ 147s, 10k token prompt @ 463 tokens/s = 22s. Hardware buy-in USD$7K / AUD$8.5K / EUR€6.8K. At typical workloads, ROI is still poor vs. current-era subsidies, but owning hardware is good for privacy/longevity/connectivity independence. Whether you actually consider Apple hardware 'owned' is a valid and thought provoking question.

    • Still gonna take 2-3 years to get DeepSeek V4.1 Flash quality at decent speeds on reasonably priced hardware.

      Hardware update cycles are 2-3 years even on the high end, so it's still a ways away before "good enough" and "local" belong in the same sentence for the average person.

      And by then, DeepSeek V6 Flash will be too cheap to meter, 5x faster, and 10x better, so... You'd still need to go out of your way.

      Most people are spending most of their time on their phones anyway. ..

      1 reply →

Not quite: not all of this needs to be in VRAM

It has a set of n-gram tables which you can stream from system RAM or even NVMe

That said it’s still quite big! I can’t fit it on my DGX Spark, though I believe you can if you have two?

  • I have access to two and will explore this the coming weeks.

    • Also give GLM 5.3 Flash a try: it’s shockingly good too in my testing, and I believe eugr has a TP=2 recipe to use for sparkrun

"1070s or 1070 TIs because GPUs have been severely overpriced for too long" ... ."

1070ti launch MSRP was $450 ish. 5070 could be had in the last year for 5xx-6xx range easily.

All things considered - (inflation being about 30%~ (guess)) between these two timelines. You are looking at 300% performance difference at a cost dollar for dollar that is cheaper then when they purchased their cards.

Might be a bit of a stretch blaming it on "severely overpriced for too long..."

Thanks for the data!

Allow a question from someone who’s only got a very vague idea of how this kind of stuff works behind the scenes: say I rent usage of this model through one of the many LLM hosting providers out there, and let‘s assume I use it extensively through something like Pi or OpenCode and vibe code away all the time, keeping the hosted model occupied as much as I can, happily burning my credits.

Does that mean that there is a hardware cluster as described by you above that is crunching away just for me?

So at FP16, I alone keep a 1,664 GiB system occupied all the time?

  • It depends hugely on what "rent usage of this model through one of the many LLM hosting providers" means. If you're asking them to host the model privately then yes, all of that 1.6T of RAM is likely in use holding weights, activations and KV cache by an inference engine that's only getting/answering requests from you alone. When you aren't actively using the model the hosting process is still active and waiting with all of that memory still wired to it.

    As background: For the most part VRAM oversubscription/paging/swapping isn't a thing in the same way that RAM for a VM often is. There are some approaches to it, but (to my knowledge) not at that sort of scale.

    There are some systemic reasons for this, but very broadly speaking the GPU vendors are building toward the highest bandwidth and lowest latency possible, and the overhead/complexity of something like protected memory modes serves neither of those priorities.

  • No, a cluster can server multiple users at the same time, providers cap the tok/s so that one cluster can run inference on multiple inputs at the same time. OpenAI with their new ultrafast mode is probably reserving the whole cluster or prioritizing requests of ultrafast users above others with a higher tok/s hence the high price and high speed. There's many other knobs providers tweak that they don't show the users, for example I doubt many providers are hosting the full FP16 version.

The article isn't just about running locally though. The author is saying it's super cheap to run the model through Opencode Go (and presumably OpenRouter etc.) Personally I'm always most excited by models I can actually run locally, but even these huge open source models open up the competitive landscape for companies to let you call models via an API or just lease compute. And they don't have to charge you to offset research, training, huge staffs of the best minds in the world, crazy PR etc. I think that's a big win for customers and buts competitive pressure on the frontier labs as well.

I think, considering the size of this model, it's closer to a Pro than a Flash on everything other than speed

I did somet math and completely gave up on the idea of trying any worthwhile local model and figured I'd rather pay the 15-30 USD per month via subscription and/or API key combos for years than buying a local setup which might go out of date very fast, if it doesn't goes kaput just out of warranty. I won't be surprised if RAM scarcity is an concerted effort to herd people towards the remote models :)

> Even so why would anyone not sleep on a model they cannot run?

Because it's an open model so providers compete on price.

I just un-retire my pair of 1080Ti for some small models development because the current GPU prices literally make me sad.

It's dangerous to go alone. Take this: [1]

I reimplemented most of the features of the Deepseek v4.1 flash paper (apart from quantization aware training which doesn't make sense because my implementation uses float32 precision anyways)

I'm currently learning how to distill reasoning traces (check my other github repositories) but I think that a locally selfhostable deepseek is possible with my mixture of experts sharding mechanism. I decided to optimize everything for CPU parallelization, with the idea that the KV cache and meta model have to run from CPU RAM anyways, so the experts can also be loaded/unloaded at runtime if needbe, to save more RAM.

My assumption is that the KV cache optimizations in combination with the CED and compressed attention features are the reason why v4.1 flash has so few hallucination problems and such a strong self-lookup/thinking behavior. But that's more a gut feeling, need to evaluate and test this more thoroughly.

Anyways, would love to see someone train this on their own datasets. Currently my pipeline is kinda optimized for parquet and zim files.

[1] https://github.com/cookiengineer/gonano

The bulk of its weights are natively MXFP4. And engram values don't need to be in vram.

He doesn't ruin the cost of memory. Advances in memory size and speed are now in full speed mode. Expect drastic increase in the upcoming years. Big factories are in the making and planned. Gigalab in the US and many others in the east. Since 2010 we have computers with 16gb as being normal. Finally we are moving into a new era where the standard will be 64gb next year and 128 in 2028. Hopefully we reach 1tb in 2030.

You can run it locally for the price of a decent car, or run it (hopefully) privately on somebody else's hardware at vast.ai or a similar provider for much less. What's not to like?

No, you won't get frontier-level intelligence on a 1070Ti. Yes, it should be illegal to do what Altman did. Since we clearly don't live in the best of all possible worlds, we need to settle, and DS4.1 Flash is a good place to do that.

For tasks that don't require vision I personally like the NVFP4 quant of GLM 5.3 from Local Inference Lab better than DS4.1F, but they are both well beyond awesome.

What are you talking about? The model is native NVFP4, why you run it at any precision higher than that?

This “blame sama for memory prices” meme is so tired.

He gave demand signal so many times years ago and was mocked for it and now we have the consequences of industry not taking him seriously.