Comment by amelius
9 hours ago
> have not been a winner-take-all runaway acceleration game where catchup is impossible
From the Mistral site:
> ML4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own datacenters in Europe.
It is pretty capital intensive!
That cluster is literally orders of magnitude smaller than the compute pools used by Anthropic or OpenAI.
For training or for inference?
They don't publish numbers, but Anthropic has a single DC with 200k+ GPUs for inference, GPT-6 Astra is said to have trained on 100k+ GPUs.
Both, especially for training. Astra and Fable were presumably trained on cluster of 100,000k GPUs, or at least a couple of 10Ks.
3,800 GPUs is nothing in the frontier side.
I’m pretty impressed that they managed to get that close to the frontier with such a small cluster!
> I’m pretty impressed that they managed to get that close to the frontier with such a small cluster
Chinese companies also managed to put together their models with relatively small clusters.
Perhaps US companies are desperately trying to brute force their way into workable models?
According to Grok thats 7-10 MW. Tiny numbers.
To put that into context, the last wave of capacity SpaceXAI added 400-450 MW.
But how much of that are they using for training versus inference? They're serving quite a large user base.
These cards are like $3k each? That's, what, $12M and you keep the hardware? Honestly doesn't seem too bad.
More like $30k each.
Oh the server chip is 10x. That makes a lot more sense.
That’s kinda very small and light for modern trillion-param LLMs.