Comment by bitexploder
7 hours ago
I am benching Flash next on a 3 bit XXS quant and it is holding just fine against published benchmarks. Using DeepSWE official harness and Pi with absolutely zero benchmaxx or harness config. Install stock Pi and running my agents in it. I am halfway through DeepSWE (it takes FOREVER, even at 125 t/s) and it is neck and neck with Opus 4.7 and Sonnet 5.
On a 3 bit quant btw.
I was skeptical but these results are simply reality now. People have figured out how to selectively quantize the tensors that matter less and shrink these models without losing quality or reasoning. This little Flash Next model just gets things done and is honestly pretty pleasant in terms of its mannerisms :)
It is so surprising to me I don't begrudge people their skepticism but these models from Alibaba represent a fundamental and irreversible shift in what local models can do. Qwen 3.8 27B and Flash Next 3.8 are simply different. But people will catch on. I am doing this on $1500 of data center leftover GPUs (V100)
How much VRAM total you using for this? I have a bunch of 3060s in a threadripper and thinking I might need to give Flash Next a try. Currently using 27B.
64GB across 2 cards. You need the MoE caching build and enough regular RAM to fit everything else to get decent performance. You should get reasonable performance if you have enough RAM. I have a 4080 that does around 40 t/s right now on a MoE caching build. It's great. Not fast, but chugs along.