Comment by peri-cl
7 hours ago
Surprisingly, the Reddit crowd are reporting 50–60 tokens/s (for the 32 GiB 5090 + 128 GiB RAM)—on par with the M5 Ultra benchmarks, despite both the PCIe bottleneck and much smaller DDR5 bandwidth,
https://old.reddit.com/r/LocalLLaMA/comments/1wl06np/qwen38f...
(Note it's a sparse MoE with only 6B active).
I have a 3 year old gaming system. RTX 4080 w/128GB of DDR5. It runs Qwen 38 Flash around 44-40 t/s with 128K context. It is on a specialized build that caches MoE experts and uses an optimized 3bit quant that basically is within a few points of the full 8 bit quant. In general, in casual benchmarking with Alibaba's endpoint I could not tell much of a difference. Overall this model is very good on long horizon agentic work. The main pain point for it is that its input processing speed is slow. Regardless, it gets meaningful work done.
I paid $500 for the RAM in Nov 2023 :)
> "I paid $500 for the RAM in Nov 2023 :)"
No wonder Warren Buffet gave up and resigned.
Right? To get comparable brand and quality DDR5, which isn’t particularly great at AI anything it is ~$2000. All you had to do was start hoarding 3090 GPU and RAM in 2023. It is unhinged.
Good to know thanks.
That's with CPU offload to a DDR5 6000 RAM though which is around $3-4k at least.
Unlike a 256GB M5 Ultra that is $10k+.
Apple product won't be the cheapest but it is a full package (CPU, RAM, VRAM/GPU, fast-storage, etc).
If you look at the pricing of a full (x86) AI workstation you'd need around the nvidia GPU, you'd approach $10k easily (and be using a ton more wattage too).