Comment by sigbottle
10 hours ago
It's interesting though that Q4 seems to be enough, is there a reason that 4 bit floats are good enough for inference?
10 hours ago
It's interesting though that Q4 seems to be enough, is there a reason that 4 bit floats are good enough for inference?
Is Qwen 3.8 at Q4 good enough?
I tried to run 3.5 27b Q4 on what local hardware i had (only 8 Gb) and i was very disappointed. 3.8 wouldn't have fit in my VRAM and i wasn't in the mood to leave it overnight at slow speeds so I didn't try.
I've run 3.8 flash next k4_xl on my Strix halo box (128GB). And in the work I have done so far it was not significantly worse than recent GPT (running default model on pro plan). Admittedly I was not doing complex work (reorganizing a jupyterbook), but I could not see significant difference in the quality of the work. It was a striking difference to Laguna s 2.1 which I had tried just before (much faster and much better quality).
At Q what? That's what I'm mostly asking about.
My little test was "generate me a single page tic tac toe game in plain javascript. computer always play O. add unbeatable minmax. have the board, a status line and a new game button'. I used both lm studio and whatever the name of their new coding assistant that supercedes lm studio is.
Qwen 3.5 Q4 went into some kind of loop where it fixed whatever was broken on the previous iteration only to have it broken some other way. (I was writing the description of the errors).
Paid $20/mo claude opus did it right the first time. Or at worst it fixed the code based on descriptions without entering a breakage loop, iForgot. I know it isn't fair because it has 1 million tokens but still, it was just tic tac toe.
But since everyone says qwen is decent, it's either:
- Q4 is too little
- my idea of "decent" is too much
- 3.8 is much better than 3.5 even at Q4
3.8 is a huge step up from 3.5, quantized or not. I do all of my programming on Qwen 3.8 27B Q4 these days.
New models are trained with 8/4bit quantization in mind. Going from "native" 8 to 4 isnt as big of a step as going from 8 to 4 if native is full bf16.
3 is the magic number, and 4 > 3.
(seriously, nobody knows why any of this works; it's just a matter of trying)