← Back to context

Comment by incognito124

9 hours ago

Qwen 3.8 flash next is way better than 27B. It's so good I dont even use claude anymore

I have not tried Flash Next yet; but 27B is a cracking, little model. It is the first small model that I, as someone with 30 years of experience, can finally say is good enough to hand off small and mid-sized tasks and expect a pretty good result.

It is also a competent tool caller when quantised to NVFP4 for use with ninfer; my own harness only reports the occasional hiccup and it is only because the model will sometimes emit tool calling tokens in its reasoning loop.

Is this true for 27b Q4_K_XL vs flash next IQ3_S? I thought under Q4 models start quickly degrading?

  • While this is generally true, it's _a little_ less true the larger the model is.

    Also, quantization techniques have improved - the I in IQ3 stands for imatrix - Importance Matrix - it is a bit more surgical in what it cuts. The result is a model where the most important weights are even Q6 or above, the least important Q2 or even below, overall it takes the space of a Q3 but with better results.

  • To add to the other comment, there's also Ridge quantisation - the majority of weights are indeed Q3_x, but the most sensitive layers are FP8.

Yeah I agree, I'm running it with Pi didn't notice much difference compared to lower tier models and the speed, of course.

  • I am running 27B with Deepseek Harness these days and somehow just by using it, without any parameter changes, the model feels even more intelligent.