Comment by spindump8930
1 day ago
The problem (and contrast with other approaches) is that mat muls requires synchronization. Arranging your networking and training structure to maximize compute and minimize communication is the main craft of ML training infra folks. In your example, yes you can compute layers on different machines (i.e. Tensor Parallelism), but you must be very careful in how you arrange it.
No comments yet
Contribute on Hacker News ↗