Comment by cmrdporcupine

2 days ago

MoE models have less active parameters in play at any particular moment, so they perform much faster on lower bandwidth memories (like Strix Halo or DGX Spark) and also use less memory generally.

Dedicated GPU memory tends to be much faster (either DDR6 or HBM). So if you can fit a whole dense model in it, you're probably better off with that.