← Back to context

Comment by drubs

1 day ago

Outside of ML metrics, you're monitoring the health of every piece of hardware in the system. You need to make sure that you have every GPU, every CPU, the PCIe buses, the networking fabric are all working without any errors. You need to ensure that you can respond as fast as possible to any possible error. One bad component can bottleneck the entire job.

https://github.com/facebookresearch/metaseq/blob/main/projec...

I really enjoyed reading the log book from the training of OPT-175B at Meta… I guess it’s all classified info but it’d be fun to read a blog post about the crazy day to day issues you run into when doing things at this scale :)

  • GPU failures are frequent enough that at a certain scale, you constantly have workers dropping out. Designing systems that can still keep training is very interesting!