Comment by abixb
2 hours ago
I have a question, and perhaps some of the AI/ML infrastructure experts here could answer: Mistral says, "ML4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own datacenters in Europe."
If a 1T params model trained on ~4k NVIDIA GB GPUs could almost match the performance of Kimi's K3 (which is on par with top closed source models of OpenAI/Anthropic) while beating/exceeding other leading SOTA models from top Chinese labs, what are we (in the US) even building these super massive data centers for? Just to churn through more backpropagation reps more quickly?
SpaceXAI's Colossus supercluster in Memphis and Colossus 2 in Memphis/Mississippi (Southaven) are supposed to run into hundreds of thousands to a million GPUs. MSFT's Fairwater GPUs are supposed to have hundreds of thousands as well. So, 3800 GB GPUs are an absolute drop in the bucket. I don't understand the strategy of hyperscalers here, especially with edge inference hardware only getting better from here on (Apple, and all).
Distillation explains some of the advances, but doesn't that mean hyperscalers have a ton of deadweight wrt GPUs sitting on their balance sheets? Will all these GPUs be used for inference once a SOTA model's training checkpoint/batch is done? It's bonkers to me.
> Will all these GPUs be used for inference once a SOTA model's training checkpoint/batch is done?
I have no real data to back this up, but that has always been my assumption.
Claude says that K3 can be assumed to have required 10-100M GPU hours. If you have 100k GPUs that would mean like 6 weeks of training. 100k GPU's can serve 3-30 trillion tokens of K3 per day. Google apparently serves ≈100 trillion per day [0].
The big labs probably want to have capacity to fairly quickly train / post train different SOTA models continuously + being able to serve peak inference demand in valuable markets (US daytime?).
[0]: https://blog.google/innovation-and-ai/sundar-pichai-io-2026
This seems to imply that training will ever be done? But yeah, I think the idea is that the appetite for thinking-on-tap will be enormous.
Even for smaller models, I think they’ve found that training an enormous, inefficient model and then distilling it internally to something much more efficient to serve is the way to go.
It's unclear how big a role distillation plays, but it may be a big one. There's also a law of diminishing returns. To get a meaningful increase in model quality you seemingly need an exponentially larger model. And most people don't think K3 is actually on par with top closed models.
Google says 1 million blackwell gpus are being delivered monthly I'm really curious if it's companies just hoarding chips/memory/servers awaiting to be deployed in data centers not ready yet for months or years, or everything built is actually deployed upon delivery. Plus google and amazon have there own chips in the mix.
Your just missing many things and so have an incorrect picture of the situation - k3 is not on par with fable or astra. Closed models remain far better than best open weights at least today. - compute is not just used for training, more and more is inf - even in training you don’t do one run, you do many. Final run is a small portion of total compute.
The world is extremely compute constrained currently, like extremely.
This is evidenced by the prices on every large-model capable device/node increasing significantly over the last year. Demand is far exceeding supply.
For inference (Anthropic and OAI are B2C on top of B2B)
To train much larger models. It is quite possible that 10T-100T models be on the horizon
> what are we (in the US) even building these super massive data centers for?
Partially to make investors think it's worth giving US companies a lot of money. Also, I think distillation is a significant part of why Chinese models perform as well as they do. I think that's completely fair play (OpenAI/Anthropic/Google/Meta stole a lot of their training data). But I would expect if US model developers stopped right now Chinese development would slow down.
US companies are clearing the path, others follow in their wake.
Today's data centers are being built for yesterday's inference need. There's a persistent cult belief that ai hasnt found a niche or that companies havent proven utility or use cases or whatever. the demand for ai (internal to hyperscaler, and external for everyone else) simply dwarfs what is available.
Both Anthropic and OpenAI have been having major load issues though. Up until yesterday, OpenAI was serving at only 30t/s per default.
You are in violent agreement with the comment you replied to.