Comment by Aurornis

6 hours ago

I have an M5 128GB. Being on the cusp of practical is a good description. It will run, but prefill and token gen are still slow relative to my consumer GPU box.

It also gets very hot. If you’ve never heard the fans on Apple Silicon really spin up, it could surprise you. Makes the full GPU setup feel quiet by comparison.

I think after the hardware market calms down the ticket is going to be a light laptop with a second dedicated inference server on the network.