All of these projects targeting low spec systems and "100 tok/s" are the same 2 bit quant without much else. Conveniently none of them include any mention of accuracy in their published numbers. 4 bit is the floor.
I am benching Flash next on a 3 bit XXS quant and it is holding just fine against published benchmarks. Using DeepSWE official harness and Pi with absolutely zero benchmaxx or harness config. Install stock Pi and running my agents in it. I am halfway through DeepSWE (it takes FOREVER, even at 125 t/s) and it is neck and neck with Opus 4.7 and Sonnet 5.
On a 3 bit quant btw.
I was skeptical but these results are simply reality now. People have figured out how to selectively quantize the tensors that matter less and shrink these models without losing quality or reasoning. This little Flash Next model just gets things done and is honestly pretty pleasant in terms of its mannerisms :)
It is so surprising to me I don't begrudge people their skepticism but these models from Alibaba represent a fundamental and irreversible shift in what local models can do. Qwen 3.8 27B and Flash Next 3.8 are simply different. But people will catch on. I am doing this on $1500 of data center leftover GPUs (V100)
How much VRAM total you using for this? I have a bunch of 3060s in a threadripper and thinking I might need to give Flash Next a try. Currently using 27B.
You can run IQ3_XXS, IQ3_S, and IQ4_XS too. I've switched to IQ3_XXS and am running at 60 t/s on Strata vs the 21 t/s I was getting in llama.cpp. Better outputs too.
Larger models can tolerate Q2 quantization surprisingly well, especially if they were trained with quantization in mind. I don't know about 3.8 125B, but for example, there are 2-bit quants of Kimi K3 that exhibit strong reasoning and maintain decent coherence at longer contexts.
All of these projects targeting low spec systems and "100 tok/s" are the same 2 bit quant without much else. Conveniently none of them include any mention of accuracy in their published numbers. 4 bit is the floor.
I am benching Flash next on a 3 bit XXS quant and it is holding just fine against published benchmarks. Using DeepSWE official harness and Pi with absolutely zero benchmaxx or harness config. Install stock Pi and running my agents in it. I am halfway through DeepSWE (it takes FOREVER, even at 125 t/s) and it is neck and neck with Opus 4.7 and Sonnet 5.
On a 3 bit quant btw.
I was skeptical but these results are simply reality now. People have figured out how to selectively quantize the tensors that matter less and shrink these models without losing quality or reasoning. This little Flash Next model just gets things done and is honestly pretty pleasant in terms of its mannerisms :)
It is so surprising to me I don't begrudge people their skepticism but these models from Alibaba represent a fundamental and irreversible shift in what local models can do. Qwen 3.8 27B and Flash Next 3.8 are simply different. But people will catch on. I am doing this on $1500 of data center leftover GPUs (V100)
How much VRAM total you using for this? I have a bunch of 3060s in a threadripper and thinking I might need to give Flash Next a try. Currently using 27B.
1 reply →
You can run a 4 bit quant with this. Personally, I switched to running IQ3_XXS and am getting better outputs than 27B and at faster speeds.
You can run IQ3_XXS, IQ3_S, and IQ4_XS too. I've switched to IQ3_XXS and am running at 60 t/s on Strata vs the 21 t/s I was getting in llama.cpp. Better outputs too.
These "revelations" are getting closer and closer to "download RAM for free" each day.
Larger models can tolerate Q2 quantization surprisingly well, especially if they were trained with quantization in mind. I don't know about 3.8 125B, but for example, there are 2-bit quants of Kimi K3 that exhibit strong reasoning and maintain decent coherence at longer contexts.