Comment by anon373839

2 hours ago

It is costly, especially right now. I don’t think you can make a case for it on cost savings!

The throughput in a single stream is about 50 tokens/sec (a bit less for prose, a bit more for code due to speculative draft acceptance rates) and about 2,000 tokens/sec for prefill. Both numbers are flat and stable as context accumulates. That’s what finally tilted me away from the Mac Studio despite its much superior memory bandwidth.

I think these numbers may improve because the model is pretty new and optimizations aren’t done.