Comment by chacham15
8 hours ago
the machines arent optimized for it is why. the main driving factor is large unified memory which makes large(r) models possible, but there isnt the gpu horsepower to back it up. essentially, it fits the corner of the market that wants large models and is ok with running them slowly which doesnt sound like it would be a large market.
For decode, memory bandwidth is the main bottleneck, so these machines will likely perform well even without a ton of GPU horsepower. Not as well as Blackwell, but I expect they will be a reasonable choice in terms of price/performance if you want to run large models with a lot of context.
The main place they are a bit behind is in the number formats they support natively. Iirc M5 doesn't have native FP8 support, so you will take a speed penalty on quants where other architectures get better acceleration.