Comment by fsiefken
7 hours ago
I wonder if a higher Qwen3.8-27b quant could beat or match these lower < 16/24/48/64G Qwen3.8-Flash Next quants given similar quality.
What speed are you willing the sacrifice to debug/program for more complex jobs faster?
Then there are also these quants; https://huggingface.co/IsValorum/Qwen3.8-35B-A3B-Distill-MLX...
Maybe, but those higher quants would need a large GPU accessible memory space - and the obvious candidates such as DGX Spark and Strix Halo don't have the bandwidth to run 27B at high quality quickly.
With Flash Next you only have ~6B active parameters so you can toss experts up into VRAM and/or run them on a CPU if you have enough RAM and bandwidth.
The benchmark indicates the IQ3_XXS quant beats 27B. I've switched to that now and am ditching 27B. Genuinely better results so far.
I'm not going to knock off my 27B-Q_6 for this. Good to to experiment though.
from my experience dense models like 27b suffer less from quantization compared to large MoEs