Comment by jboss10

8 days ago

Most of the time, the speed of these models are constrained by memory bandwidth. GPUs normally have much more memory bandwidth.

I'd expect the memory bandwidth to be the same for the CPU and GPU under a unified memory architecture like Apple silicon uses?