Comment by nacs
7 hours ago
That's a dense model. Of course it will do worse.
Now try running that Qwen 3.8 Next model on the 5090 and tell me what TPS you get (hint: it's near 0 since it doesnt fit the 32GB VRAM on 5090 vs the 256 in OPs M5).
7 hours ago
That's a dense model. Of course it will do worse.
Now try running that Qwen 3.8 Next model on the 5090 and tell me what TPS you get (hint: it's near 0 since it doesnt fit the 32GB VRAM on 5090 vs the 256 in OPs M5).
Surprisingly, the Reddit crowd are reporting 50–60 tokens/s (for the 32 GiB 5090 + 128 GiB RAM)—on par with the M5 Ultra benchmarks, despite both the PCIe bottleneck and much smaller DDR5 bandwidth,
https://old.reddit.com/r/LocalLLaMA/comments/1wl06np/qwen38f...
(Note it's a sparse MoE with only 6B active).
I have a 3 year old gaming system. RTX 4080 w/128GB of DDR5. It runs Qwen 38 Flash around 44-40 t/s with 128K context. It is on a specialized build that caches MoE experts and uses an optimized 3bit quant that basically is within a few points of the full 8 bit quant. In general, in casual benchmarking with Alibaba's endpoint I could not tell much of a difference. Overall this model is very good on long horizon agentic work. The main pain point for it is that its input processing speed is slow. Regardless, it gets meaningful work done.
I paid $500 for the RAM in Nov 2023 :)
> "I paid $500 for the RAM in Nov 2023 :)"
No wonder Warren Buffet gave up and resigned.
1 reply →
Good to know thanks.
That's with CPU offload to a DDR5 6000 RAM though which is around $3-4k at least.
Unlike a 256GB M5 Ultra that is $10k+.
1 reply →