> So given that, which one of these is true about you?
Well, if those are the only two options you can come up with it's pretty clear that this isn't about me or what I am, you have a false model of reality.
> Running very large models on Mac is unusable at 10 tok/sec.
There are plenty of examples of models running at well over 10 tok/sec that aren't viable on the 3090. In fact such examples are found in the review in the OP. Did you not read the article?
I think you're projecting pretty hard with the two options you've listed. Go touch some grass, you seem overly frustrated that reality doesn't meet your expectations.
Since you clearly don't use local llms, allow me to educate you - anything under 100 tok/sec is USELESS. When you are coding, the idea is that you want to have a system that can generate files fast, hopefully correct on the first try. Cloud models do this. Local models, by nature of having less parameters and more quantization, often require more guidance and repeated inference to get it right. The antigenic harnesses that people set up around local llms leverage this.
Looking at the article, which you clearly didn't read,the m5 ultra runs Qwen3.8, which fits on one GPU conveniently, at ~20 tok/sec. This is a fucking joke. It will take roughly a minute to generate one code file. Congrats if you want privacy I guess, but for straight up coding, you are better just using cloud models.
Meanwhile, I have an $800 mini PC, $200 Occulink gpu dock, a $2000 3090 and a $300 power supply, and I can run Qwen at over 100 tok/sec prefill, not to mention insanely quicker during inference.
So its pointless to spend Mac M5 Ultra prices on Apple shit when they can have something much faster for cheaper
The whole thing of "well I can run bigger models that don't fit on a GPU" is either paid Apple advertising, or you are just an igorant fanboy.
[flagged]
> So given that, which one of these is true about you?
Well, if those are the only two options you can come up with it's pretty clear that this isn't about me or what I am, you have a false model of reality.
> Running very large models on Mac is unusable at 10 tok/sec.
There are plenty of examples of models running at well over 10 tok/sec that aren't viable on the 3090. In fact such examples are found in the review in the OP. Did you not read the article?
I think you're projecting pretty hard with the two options you've listed. Go touch some grass, you seem overly frustrated that reality doesn't meet your expectations.
Since you clearly don't use local llms, allow me to educate you - anything under 100 tok/sec is USELESS. When you are coding, the idea is that you want to have a system that can generate files fast, hopefully correct on the first try. Cloud models do this. Local models, by nature of having less parameters and more quantization, often require more guidance and repeated inference to get it right. The antigenic harnesses that people set up around local llms leverage this.
Looking at the article, which you clearly didn't read,the m5 ultra runs Qwen3.8, which fits on one GPU conveniently, at ~20 tok/sec. This is a fucking joke. It will take roughly a minute to generate one code file. Congrats if you want privacy I guess, but for straight up coding, you are better just using cloud models.
Meanwhile, I have an $800 mini PC, $200 Occulink gpu dock, a $2000 3090 and a $300 power supply, and I can run Qwen at over 100 tok/sec prefill, not to mention insanely quicker during inference. So its pointless to spend Mac M5 Ultra prices on Apple shit when they can have something much faster for cheaper
The whole thing of "well I can run bigger models that don't fit on a GPU" is either paid Apple advertising, or you are just an igorant fanboy.
So I ask you again, which one are you?
2 replies →