Nobody knows what a used GPU cluster is worth

8 days ago (ciphertalk.substack.com)

"Secondary market data shows moderately-used 2 to 3 year old GPUs trading at 50% to 70% of new pricing under normal conditions."

That's for good NVidia H100 units.[1] There's a shortage of those. That seems to be the price after removal, cleaning, testing and refurbishing. Raw units removed from a shutdown will not be as valuable.

H100 units are available on eBay, but multiple sellers are using the same picture of a new unit in its original packaging, a bad sign.[2] Some even have pictures with the logos of a competitor.

[1] https://introl.com/blog/secondary-gpu-markets-buying-selling...

[2] https://www.ebay.com/shop/nvidia-h100-gpu?_nkw=nvidia+h100+g...

  • re: your last paragraph, there's probably only about 30 to 40 (at max) reputable relatively high volume dealers of used/refurb ex datacenter server equipment dealers on ebay that are located in the US48 states. It would be very risky in my opinion to buy a used GPU or multiples of GPU from some rando who has 14 feedback.

    If you search ebay for server equipment like a Dell R840 with 768GB RAM, the same sort of dealers who are selling that and have thousands of feedback (at 98.5% of greater rating) are the ones I would consider much less risk.

  • Who is buying these at 70% of new pricing given the sky high likelihood of them being shot? Maybe it's safe to buy from small labs that went under quickly, but I can't imagine a cluster that has been operating near its thermal limits for a couple years fetching that kind of resale.

    • Why would they be shot? Unlike the consumer cards that are basically factory overclocked to look good on benchmarks, the datacenter GPUs are designed to run at full tilt 24/7 and survive for years.

      4 replies →

    • It was true of crypto GPUs too, although mostly people picking them up for gaming. Always seems high to me too but if you can get any guarantee of them not being on fire when they were pulled the bathtub curve keeps you pretty safe, thermal limits are limits for a reason.

      9 replies →

I half expect Nvidia to have buyback contacts like Ferrari with the larger customers to prevent a price crash when they all upgrade and to keep them scarce.

I hope not though, perhaps I can pick up a H100 in a few years if they get sold on the open market.

  • I worked at an org that had a substantial on-prem GPU datacenter. We transitioned to <Big Cloud Provider> with a substantial negotiated discount rate, with part of the contract being we would sell them all of our hardware and not purchase any more.

  • There's going to be a golden age of GPGPU compute in the next few years once A100/H100 are fully obsolete for running frontier models efficiently and the price plummets

    It will be perfect for stuff like GPU-accelerated query engines, "classical ML" and every other CPU-based workload that could conceivably be offloaded to GPU

    • There's nothing stopping you doing this now.

      You can get used 16GB P100s on AliExpress for ~$100 if you want obsolete GPUs. Allegedly new AMD BC 250s are only slightly more.

      I've looked at this some but I already have a GTX1070 which is only supported upto CUDA 11.9.

      That's precludes some interesting modern optimizations out of the box. I've spend a lot of LLM tokens backporting some things, but I'm really not sure the hassle is worth it.

      New hardware is just better. I think in maybe 5 years when supply and demand are back in equilibrium we are going to have some killer technology for decent prices, and 15yo H100s won't look attractive.

    • GPGPU? General Purpose GPU? If embarrassingly parallel CPU algorithms weren't offloaded to the GPU previously, why would the A100/H100 price drop make a difference? We had cheap GPU in the past and we still left plenty of performance on the table with CPU programs because they were easier to build.

      Is the idea that previously maintaining GPU programs was expensive whereas now AI makes it cheap? If so, I could buy that line of reasoning.

      Maybe relatedly, I expect (hope) the hardware manufacturers will ramp up supply in the meanwhile which would also put downward pressure on GPUs. Right now though this hardware crunch is making me sad, not even because of GPUs but also because of general memory / disk.

      4 replies →

  • The only reason a datacenter would ditch their H100 is if it becomes uneconomical to run them, with newer silicon providing much more power efficiency. When that happens they'll look like a used V100 looks today: horribly inefficient, lacking modern data types and engines, requiring screaming server fans with weird adapters to not melt, way beyond end of life in terms of cuda support. Almost completely damn useless unless you really have no other alternative.

  • That should be illegal. Sounds like a very fraudulent business tactic.

    • The hyperscalers signing these contracts have decent legal departments. Think about Oracle for example - I'm pretty sure they know every trick there is about beneficial contract drafting.

      I don't think they need some special protection against this kind of contract.

    • It’s common for car companies when they enter a new market. It removes uncertainty from the second hand market.

      By doing that, you know upfront what the value of your used hardware will be at the time you decommission it. It removes a lot of the risk for buyers in a volatile market.

    • The Grift Economy places all legalities on the marks and their inability to form legal fights.

  • Wouldn't that be crazy - a hobbyist market for H100s?

    Maybe someone could start a business buying up and rehousing these.

    • They're pretty specific to the datacenter use case with no outputs and they need to be cooled externally, principally through the very loud high speed fans used in data centers. I suppose you could strap a fan to one and put it in a normal case or maybe make a dedicated after market cooler (like the water blocks made for water cooling cases).

      11 replies →

    • H100s (and the PCIe converter card you need) are available on eBay.

    • The GPU alone has a TDP of 700W, together with everything else (CPU, RAM, storage, fans) you're looking at 1500W+. Depending on the country, that may be enough to saturate your home's electricity uplink...

      10 replies →

  • isn't Ferrari the brand that requires any purchaser to be an existing owner? I could see NVIDIA going for something like that.

The relatively slow depreciation of GPU value is an artifact of supply constraints. If you run fp4 inference and could choose freely between Hopper and a Rubin, the performance per watt would make the Hopper unattractive even if you paid zero for the hardware and only for the power.

You can't get the Rubin, or even the Blackwell, so you will pay for the H100 but this won't last if fabs ramp up capacity.

  • Not to mention, physical limits to lithography are slowing down significantly... so tech will continue to evolve more slowly... it'll never be the jump from 1080-1990 again, for example, even though 1990-2000 was pretty close, 2000-2010 much slower and since 2010 slower still.

    What's as or more weird is how much hardware is backordered, and how much live hardware is allocated, but waiting on facilities for operation. And how many facilities are years behind at this point already... all on various credit and dept swaps between all the involved companies... it's not just a balloon, it's a house of cards balanced on a balloon.

    • The infamous attention is all you need paper has this: "The Transformer ... reach a new state of the art in translation quality after being trained for as little as twelve hours on eight P100 GPUs.", the p100 is now quite a bit under 100$ and that's from a time where nvidia was valued 100x less than it is today. Lets wait and see

    • > since 2010 slower still

      The increase in PFLOPS/dollar has continued accelerating, a lot from process, but also a lot by simplifying the architecture- if you had placed an H100 worth of transistors on a CPU-like architecture, you wouldn’t reach the same peak performances.

    • supposedly, the next step is into fiber & optics.

      but I generally agree, people put a lot of faith in the exponential leaps vs the exponential space.

      You tell them we're not living on mars any time soon and they'll bring up christopher columbus.

      2 replies →

  • Are fabs going to ramp up capacity? Wouldn't it make more business sense for them to just not ramp up capacity and enjoy the higher prices?

    • Yes they are going to ramp up capacity. It would only make business sense to do nothing if all your competitors were also doing nothing, which would probably require some level of illegal collision.

      If everyone ramps up then in the best case everyone has the same sized slice of a bigger pie. So in theory it makes business sense. But the more realistic possibility is that you end up with oversupply, crash the market and everyone loses. This is what normally seems to happen with DRAM.

  • I wonder what shovels were worth after the Gold Rush faltered. Blacksmiths probably had all the scrap iron they could ever care for.

    • Nah, "selling shovels" is mostly a metaphor, the amount of iron that went into mining equipment was insignificant, beyond some local demand peaks. The majority of the business was consumables.

      A key thing to understand about the gold rush is that it was not a major economic event, or at least nowhere near as big as the participants thought it would be, hence the tradegy.

      The AI gold rush is different in that there actually is a mountain of "shovels" large enough to flood the global market quite severely.

> The job of an operations team is to keep all of this in steady state. They know which racks run hot in summer, which cooling loops have been flaky since the last firmware update, which jobs to re-route when a node degrades but has not failed yet. None of that knowledge is written down. It lives in the team.

Hmmn. All of this information should live in the monitoring system, in which case any frontier model will be able to get to grips with it in short order. It feels like the author doesn't really fully understand the changes brought about by the systems they are writing about.

One data point: 512GB Mac Studios are selling on ebay for double what they were selling for new at the beginning of the year.

I don't get how this is unique to GPU clusters. As a general rule, underwriters are not qualified to operate and maintain the assets they underwrite loans for. That's why houses, cars, equipment, ... go at auction at a fraction of their value. And why lenders have insurance.

  • > I don't get how this is unique to GPU clusters

    It isn't, and you're right. This is just a long-winded article by someone who thinks they've come across a deep, crucial insight.

    I witnessed a bank foreclosure stemming from large, unpaid loans to a lumber mill. The bank absolutely didn't want to take possession of the operation, but had no choice when the founder decided to call it a day.

    The bank had no idea what to do with finished lumber sitting in the drying ovens, let alone the entirety of mill infrastructure itself. After struggling to find a buyer, they hired the founder as a consultant to handle liquidation. The same would happen with a repossessed datacentre.

  • Or... auctions are a terrible way to buy big expensive risky purchases.

    Auction prices account for that risk, you don't get to do all of the verification you do for normal purchases.

All indications are there will be a lot of repossessed GPUs appearing on the market before too long. Likely to be messy for a while but will open up a lot of possibilities when it’s easy to get your hands on some secondhand GPUs.

  • This sounds like weapons market after the crash of soviet republic. Never has been a better time to get a functioning tank or parts of nukes.

Very difficult to sense check that substack post without access to the TLB credit agreement.

There are likely management service agreements from xAI proper -> SPV to cover precisely what the author talks about. Clearly, xAI could play games but without seeing the docs (which are not public), it's very difficult.

This article's basic point is right though. On the other hand, the LTV of this deal was approx 50% debt-financed (not too high; very much depends on the "V"). At 12.5%, it's not as if its being priced as a high quality asset.

Overall, substack post was too bearish. The wider point is that there's a lot of froth tied to what has now become systemically opaque - namely the circular deal flow that every hyperscaler, nvidia, neoclouds and friends are now engaged in. When the proverbial hits the fan, that stuff will be difficult to price and find few willing buyers with the competence to underwrite.

The systemic issues are the bigger concern than one specific deal imo.

Should I be worried that my plant is owned by Apollo?

No, because I am retiring next month!

I'm wondering if all those GPUs ever end up on the second hand market for us to buy. Or if they will get refurbished multiple times and get used in second tier data centers until the chips die.

I like Meg a lot a human, but Meg is all doom and gloom. Every single post she makes is about how GPUs fail [0] and now she's onto how financing is a big thing just waiting to crash and "nobody knows what a used GPU cluster is worth"...

Actually, we do, people offer them to me all the time. A used box of MI300x is $257k. "There is no GPU futures market"... actually there are a few of them that people have pitched to me.

This article is a lot of words from someone who isn't actually buying or deploying compute. My point is... take it all with a grain of salt.

[0] https://x.com/meggmcnulty/status/2040851080066859386

  • Somehow at any given moment every possible topic to discuss on HN belongs to either the set of "it's amazing and nobody can say anything bad about it" or the set of "this thing sucks and nobody is allowed to say anything positive about it" and I never have ANY idea which one any given topic will be in on any given day.

    • latchkey is being restrained. He runs a data center filled with AMD GPUs. He's got a lot more insight to the business of it than the post does.

      4 replies →

    • latchkey is being restrained. He runs a data center filled with AMD GPUs. He's got a lot more insight to the real, lived experience of such a business than the post seems to have.

  • It’s always the people who know the least about GPU economics who want to doom post. Until I can reliably get on demand A100s in large quantities for less than 2.00 an hour there’s no AI bubble.

    Most of the people who shit talk GPUs don’t even understand why it relatively speaking doesn’t matter that much if I’m on ampere or Vera Rubin for full precision models. These same people are wildly misinforming investors.

    To quote the ADATA ceo, “don’t talk about an AI bubble until 2040”

Isn't it worth exactly what you can sell it for. a few ways to do this.

slow and awkward, best market match: The auction. sell to highest bidder.

faster and more customer friendly but poor market match until a lot of units sold: The store. guess price, adjust up or down to reach sell frequency desired.

fast and good market match but takes a knowledgeable customer base: The reverse auction. Start with price too high lower it over time until it sells.

  • >adjust up or down to reach sell frequency desired.

    >Start with price too high lower it over time until it sells.

    These are the same strategy.

This is a good article. I've been curious about how this is going to play out. A couple of data points:

1. An enthusiast had a project to get a V100 working on his PC [1]. This was a ~$10k GPU 10 years ago. It's now sold for scrap;

2. The A100 came out in 2020 and cannot run a large model like DeepSeek v4 Pro. It can run Flash. You need a 16xH100 cluster to run Pro and that's a ~4 year old GPU and AFAICT 8xB100 or 4xB200;

3. We're about to roll out R100/R200s.

I'm surprised that NVidia is moving to a 1 year product cycle (per this article) because the big question I've had is what's that going to do to existing investments in GPUs. Why? Because if 4xR100 can do the work of 32xH100 then that's a massive advantage in performance-per-Watt, which I think is going to be the only metric that ends up mattering.

In addition to raw power, new capabilities are developed and come online. For example, certain smaller, more efficient quantization methods just didn't exist on older hardware.

Oh, another thought from this: a 9% annual failure rate just goes to show you how ridiculous the idea of orbital data centers really is. Orbital DCs were always just a pump-and-dump scheme for SpaceX's IPO.

Currently it gets expensive to run models larger than ~31B locally. You start to need some pretty expensive hardware. That's going to change. I don't expect we'll be running 1T+ models on a Macbook Pro within 5 years (at reasonable inference rates) but I think people today will be shocked at what's being run locally in 5 years and that'll easily be 100-200B+ models.

[1]: https://www.hackster.io/news/hacking-a-server-grade-nvidia-g...

The GPU cluster's RAM is probably worth more than the flops these days.

A relevant article I found from an industry insider (which could indicate bias but also relevance): https://www.whitefiber.com/blog/understanding-gpu-lifecycle

Which has this anecdotal data point:

>... when I left Paperspace in mid-2024, our M4000 GPUs, nine-year-old GPUs, were still consistently utilized at near-total capacity. That’s not a typo. Nine. Years. Old. Still booked, still working, still generating revenue.

Also, I won't claim to understand accounting, but in general it seems it is advantageous to accelerate depreciation schedules for high CapEx industries because they lower taxes: https://leyton.com/us/insights/articles/what-is-accelerated-...

As such you'd assume these cloud providers to want faster depreciation of their GPU assets rather than slower? I suppose in this case they do have an incentive to show bigger revenue numbers, but there seems to be a trade-off here that is not being discussed?

Is the 30-50% of face value realistic for "Liquidation Value"? If one organization has to liquidate, sure. But if the bubble pops and many groups have to liquidate at once? Owners will be lucky if they can dodge the recycling fees.

Maybe this article has something useful to say but the painfully LLM-generated prose is too distracting to make it evident.

  • How do you know? I feel that it's just poorly edited.

    In my experience, AI is easier to read than this was.

    • Some sentences feel pretty AI like:

      > These are not catastrophic events. They are the steady state.

      > There is no GPU futures market, no standardized residual value curve, and no way to lock in a forward rental rate. The premium is is the price of underwriting in the dark.

      The headings are also AI like, a lot of essays before usually did not have titled sections but now they do and they all feel like these.

      In addition the diagrams themselves look pretty AI generated.

      1 reply →

Nothing, because those "GPU's" are special proprietary hardware and are not what most people are capable of plugging in into anything.

Essentially nothing, fractions of a penny on a dollar, because players who could afford paying real money won't risk the crusty old hardware - so you're limited to buyers who still need a massive cluster but don't have AI infinite money glitch enabled

  • Are failing GPUs really a big issue? I don't know how they behave when they die, but if it can be detected quickly, the affected nodes can just be removed from the pool.

    • I think there’s a bit more nuance and complexity here.

      I built and operated an application that used 1-2000 L4 GPUs in production for a couple of years. Long-term GCP reservations running in a GKE cluster. At that scale we had a few GPU failures per week, and once saw 3 in a day. Nearly every failure required an engineer to manually intervene to get rid of the bad node, and then file a ticket with GCP support as they requested.

      NVIDIA GPUs throw an “XID” code when they fail, which can be seen from serial port logs for GCP compute nodes (not through k8s!). If you’re lucky, it fires right as your application starts to fail, but there’s often a delay of several minutes. Even when you get an XID, by default GKE only responds to one or a few of them. You can expand the list via configuration, but the reconciliation loop is so slow that might still take 10-15 minutes during which a pod is puking errors and someone might be getting paged.

      They’re working on it, and we never saw a single XID after we migrated to H100s, so the situation is improving. I imagine other clouds are even worse, though iirc Azure was leading some effort to improve k8s node problem detector to include accelerator problems so maybe they have a better story.

      Training workloads are rather more sensitive to a node failure given that many modern training runs (including SFT etc) need multiple nodes where topology matters, and they might not have another e.g. 8xH100 box in the right place when one fails.

Would it make economic sense to strip it down and sell for parts? That’s how it’s done now for older data centers, where the obsolete equipment is sent off to China, stripped down for parts, and sold on the secondary market. I’ve picked up several older but still useful RAID hardware cards off eBay this way.

I’m mainly interested in getting some DDR4/5 and RTX5090s on the cheap :).

Given the price for a rack of bc-250 after the crypto hype cycle, the expected value of the hardware will be around 5% to 10% of the original retail price.

Without other market influences, that is a >90% expected discount when the over-provisioned market must inevitably self-correct.

If the Market follows what Samsung/SK Hynix did to the South Korean exchange this week, than the "AI" bubble will hit harder than the dot com crash.

I like the Shrek Movie correlation theory, as they always happen just before Debt-backed investors get hit hard... And the new film is due out in 2027. =3

  • > If the Market follows what Samsung/SK Hynix did to the South Korean exchange this week, than the "AI" bubble will hit harder than the dot com crash

    Can you tell us more about this? Or some link

    • Samsung and SK Hynix together account for around 60% of the Kospi's (SK stock exchange) market capitalization.

      Over the past few weeks Kospi index has tumbled 25% since its June peak, resulting in a $1 trillion wipeout and its chipmaker duo have both lost at least 30% of their value. There have been days of near 10% plunges followed by sharp rebounds driven entirely by shifting confidence in whether AI spending is sustainable.

so you're telling me you can't use 1 or 2% percent of 5 billions dollars to rebuild a team that runs GPU clusters for like a couple of years ?! and the guy that borrowed billions from these banks would want to mess up that relationship for what? I mean he is a stupid narcissist but not to that degree. This whole article makes no sense to me.

GPU clusters have, largely, no actual value.

If anything, you might have to pay to have them disposed of, they don't really have any meaningful used eBay market outside of the randos that want to do high end extreme local inference in their basement.

Also, as for RAMmageddon, the inference SBCs that all of the AI bros bought don't have DIMMs, they're not even the right chip: its all GDDR and LPDDR. The only DDR DIMMs being consumed are for regular non-inference machines that help run the business and service infrastructure behind the scenes.

> Silent data corruption (SDC) is the most expensive, where a faulty GPU produces wrong answers without crashing anything, which means a multi-day training run can complete normally and the resulting model weights are quietly poisoned.

Anyone know how these get caught ultimately?

  • Usually silent data corruption isn't caught. If a file fails to open you might realize it's been corrupted.

[flagged]

  • That still doesn't tell you what GPUs will be worth in 5 years, because it depends upon what inference demand is like and what the alternatives are. What will an hour of NVL72 be worth in 2029?

    (And other things, that we know partially but not fully-- like what failure rate for current generation parts will be under this loading).

    So we have big uncertainties about the revenue, moderate uncertainty about the proportion of the asset that will survive, and some uncertainty about what operating costs will be. It's difficult to turn this into a residual value.

    Finally, the whole "operating the big facility" thing is not likely to be plug-and-play for a new technical team following a default. How much outage/disruption ensues?

If nobody knows what a used GPU cluster is worth it means nobody is doing anything important with GPUs - time to short everything. Do you believe it?