Comment by syntaxing
3 hours ago
All these new models are such tease for us folks with 128GB of shared memory. Buying another unit now to expand to 256GB is a mortgage payment but it’s getting tempting…
3 hours ago
All these new models are such tease for us folks with 128GB of shared memory. Buying another unit now to expand to 256GB is a mortgage payment but it’s getting tempting…
https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B is an option
That’s for toy GPUs, like the 5090.
there are many tasks (increasingly more each day) where small models are more than enough
Is there a gamechanger around the corner to reduce DRAM requirements?
You could always stream from SSD storage. Especially effective if you get a cheap old-gen HEDT with lots of PCIe slots to add NVMe storage to and reasonable overall PCIe bandwidth.
That nearly certainly boots you to secs-per-tok land (as opposed to tok/s). Plausible if you are willing to wait hours to days for responses for simple testing, but not (debatably) "usable".
n-gram per-layer embeddings[1][2] might be it.
[1] https://sebastianraschka.com/llm-architecture-gallery/per-la...
[2]: See DS 4.1-Flash and Qwen-3.8-Next.
this is to offload VRAM to DRAM (for GP comment), and makes no difference for URAM
18 replies →