Comment by 0xbadcafebee
9 hours ago
Lol, sure, if you quant it to hell (Q2) it'll go real fast...
They even link to a Q1 quant (Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF) with half the experts ripped out. The idea is it'll go much faster and supposedly benches to not-terrible results. But the problem is you can't rely on it for real world long-horizon coding because that's where reasoning comes in, which is why you want the other layers.
It turns out there's still no free lunch. Either get enough VRAM for a Q4, or use a much smaller model. Lobotomizing a larger model just to say you can run it fast isn't useful.
You're right but for simple use cases its useful. Someone pointed to ds4 + qwen3.8 with Q4_K will try that out.
It's interesting though that Q4 seems to be enough, is there a reason that 4 bit floats are good enough for inference?
Is Qwen 3.8 at Q4 good enough?
I tried to run 3.5 27b Q4 on what local hardware i had (only 8 Gb) and i was very disappointed. 3.8 wouldn't have fit in my VRAM and i wasn't in the mood to leave it overnight at slow speeds so I didn't try.
I've run 3.8 flash next k4_xl on my Strix halo box (128GB). And in the work I have done so far it was not significantly worse than recent GPT (running default model on pro plan). Admittedly I was not doing complex work (reorganizing a jupyterbook), but I could not see significant difference in the quality of the work. It was a striking difference to Laguna s 2.1 which I had tried just before (much faster and much better quality).
1 reply →
3.8 is a huge step up from 3.5, quantized or not. I do all of my programming on Qwen 3.8 27B Q4 these days.
New models are trained with 8/4bit quantization in mind. Going from "native" 8 to 4 isnt as big of a step as going from 8 to 4 if native is full bf16.
3 is the magic number, and 4 > 3.
(seriously, nobody knows why any of this works; it's just a matter of trying)