← Back to context

Comment by htrp

3 hours ago

>The number traces to Meta’s Llama 3 technical report, which documented 419 unforeseen disruptions across 16,384 H100s over 54 days of training, of which 148 were GPU failures and 72 were HBM3 memory failures.

From an annualized number on the llama 3 training report. would be interesting to see if we have a better idea given that we're already on rubin.