Comment by quanto
10 hours ago
Is your Qwen3.6 locally run on your laptop? What kind of tokens/s are you getting from your laptop GPU?
10 hours ago
Is your Qwen3.6 locally run on your laptop? What kind of tokens/s are you getting from your laptop GPU?
Qwen is running on my Mac Studio, an M1 Ultra 64gb. My harness (oh-my-pi) on my laptop is configured to use the models hosted on my local network, since it's just a MacBook Air 16gb and probably incapable of running anything useful itself.
I get about 45-55 tokens per second using Qwen with this setup. I could probably squeeze out more if I messed around with the settings, but I'm mostly using oMLX's defaults for the model.
I have Qwen3.6 35B-A3B on my laptop and it does 60 tokens/s
Could you share the specs of your laptop?
M5 Pro/64 GB RAM