Comment by andy_ppp

3 hours ago

Tokens per second is almost entirely memory bandwidth at inference time, training obviously needs more compute but you can add more chips for that.