Comment by solarkraft
12 hours ago
So far I don’t regret buying an M1 Max device with 32Gb of RAM. The models available for it keep getting better (running just about okay for interactive use) and 400 GB/s of bandwidth is still considered a lot.
The models are currently improving much faster than the hardware and this doesn’t seem to have plateaued yet.
Cool! I'm thinking about a local set up. What's your usual tokens/second rate?
Not OP, but I’m running local models on a M1 Max as well with 64GB RAM.
It varies by model, but I’m getting 50-60 t/s with Qwen 3.6 35B and Qwen 3 coder 30B.
I’ve also used Qwen 3.8 27B but I get 10t/s on it.
It’s useable in some use cases, but I rely mostly on my $20 Claude subscription.
That's so cool. I wonder if the regular M5 can run those models too.