← Back to context

Comment by christina97

5 hours ago

Why would they be shot? Unlike the consumer cards that are basically factory overclocked to look good on benchmarks, the datacenter GPUs are designed to run at full tilt 24/7 and survive for years.

According to the article these cards have a 9% annual failure rate.

  • >The number traces to Meta’s Llama 3 technical report, which documented 419 unforeseen disruptions across 16,384 H100s over 54 days of training, of which 148 were GPU failures and 72 were HBM3 memory failures.

    From an annualized number on the llama 3 training report. would be interesting to see if we have a better idea given that we're already on rubin.

  • It depends entirely on the time-to-failure distribution though whether used cards are a good deal or not. Often this kind of hardware has a bathtub shaped hazard rate, actually getting burned in cards may mean you get the weeded out solid specimens, and forgo the lemons.

  • Most hardware has an exponential (memoryless) failure distribution in practice. The bathtub curve is a myth.

Given the decades of consumer and enterprise GPUs often being identical or near identical hardware, it'd be interesting to see if there's any evidence of this actually being true.