← Back to context

Comment by zozbot234

5 hours ago

First of all, a x090 series card is "high end hardware" in its own right these days. Secondly, I'd think you'd probably get more interesting results running a MoE model in CPU-MoE mode, i.e. with the shared parameters residing on GPU and sparse experts on CPU plus SSD offload. Yes it will be slower, but small dense models are just a dead end and not that interesting. (Note that prefill would still be sped up in this setting; the CPU/GPU layer split in llama.cpp and the like applies to decode, but even a "0 graphic layers" setup does accelerate prefill.)