Comment by a11r
7 hours ago
We recently moved from 27B to Flash Next. The quality is superior for coding. Our workload is primarily well-defined coding tasks that need to be attempted a few times before the model gets it just right. FlashNext is also better at finding issues in generated code than Gemini 3.8 Flash.
I'm on m1 max 64gb and went from qwen3.8-27B back to qwen3.6-a35b. Is flash next the move? I went from usable say 40tk/s qwen3.6 to unusable, like 11 with 3.8 and not impressed with the replies for the time sacrifice. pi (omp) and omlx but not with the recent 3.8 patch.
I've been waiting for a 35b of 3.8, I don't really know what the other versions are about. I'm on 5g so juggling 40gb of model files sucks. And honestly I'm sick of tweaking this stuff for no, very little, or break-it level improvements. Qwen3.6-a35b has been solid for work, just don't give it freedom to wipe your data.
Exact same scenario here
I’ve heard a quantised version of flash next can fit in ~50 gb of vram (which needs a system level flag set to go over 48gb)
But the m1 cpu is itself a bottleneck on prefill compared to say an m5, there’s no real getting around it. And the 400mb/s bandwidth starts to hurt without MOE
Hoping these model optimisations can see us through to 2028 because for everything other than LLMs this hardware is still over specced and working incredibly well