Comment by ilaksh
18 days ago
Dumb question, but when the Nebius capacity dashboard says they have around 3 non-preemptible B200s available, does that mean _total_, or is it just how many I myself might be able to rent on demand?
One aspect of the profitability might be the utilization and the pricing a few years down the line for slightly older hardware. Already now it seems like the increased processing you get from newer devices versus the cost difference makes something like an H100 or even A100 significantly less desirable than newer more powerful ones. As an individual, I am happy to be able to get an H200 on demand, but the B200 or B300 can do so much more work with optimized software and models for only modestly more cost that if those become available then from a business perspective you really have to prefer that if you can keep it occupied.
Then with Vera Rubin being like 3 times more effective or whatever, that adds a new layer of gradual obsolescence. So the question is can they keep the pricing up on the older ones a few years down the line enough to fill out the end of those expected payback periods.
The real boogeyman for a neocloud that has heavily invested in expensive Nvidia hardware might be a variation of that beyond Nvidia with startups that have even more dramatic efficiency increases pushing the leading edge even further. For example, if companies like Mythic AI and d-Matrix could somehow rapidly rapidly scale, that would push prices down for all of Nvidia hardware that is significantly less efficient.
I guess so far it doesn't look like any startups with really big efficiency breakthroughs are even close to being able to scale like Nvidia though, especially with the manufacturing and power crunch. But I suspect some of that is because of favoritism and strong arming protecting investments rather than a free and fair ecosystem.
> So the question is can they keep the pricing up on the older ones a few years down the line
They don't expect to keep the prices flat over time, and everyone involved will have planned for this. Prices are highest when they're the newest and greatest (part of why it's valuable for neoclouds to be first in line for new models), and drop year by year as newer GPU models can do equivalent work at lower cost.
You can see a pretty cool dataset of this at [1]; H100 prices where $3/hr in 2023, and dropped linear-ish to $1.75/hr by 2025. And also the notable exception that prices are up this year due to shortage.
[1] https://semianalysis.com/gpu-pricing-index/
Why aren’t you talking about how A100 and H100 prices have again spiked and in some cases are higher than 2023! (This chart you posted isn’t fully accurate but even it admits that a100 prices are basically flat since 2024)
Micheal berry doesn’t know shit about GPU pricing or depreciation schedules. A100 demand is very high and easily 2 dollars an hour for reserved right now.
B200 and Vera reubin don’t help much if you don’t benefit from quantization, and that’s exactly my situation and many other AI research orgs situation.
A100s are going to continue making money per hour until 2030. Mark my words.
hm. I should have written that more precisely, but I thought it was implied/obvious that it would go down to some degree or another. I didn't mean it would literally remain flat. it's a question of how much they drop though, and I don't think they know for sure, because there are new technologies that are really just being held back by manufacturing scale and anti-competition which otherwise could cause larger than anticipated pricing drops for older hardware. like.. how could you read what I wrote in that comment and conclude that I needed you to explain that the prices would drop?