Comment by mmaunder
8 hours ago
More great work on local model but you’re still losing a lot. Down to 2 bit quantization and the coder model throws away half the MoE experts. In a world where anything is better than nothing, this is a net win. But we have a way to go still.
The real problem is the DRAM mafia and artificial scarcity. One of my notebooks is almost 3 years old, effectively similar spec now - same price (a bit higher actually). My desktop PC built around march/april 2023 (4090, 64gb ram, 7900x3d) is now pretty much still the top dog out there due to the gpu and fast ram insanity and if I wanted to sell it today, I'd get more money for it now used and over 3 years old that when I bought it!
We should be having 64/72+ GB video cards by now. 128GB+ system ram prosumer laptops and 256GB+ system ram prosumer/gamer desktops. But it all went to shit and it will require some brutal datacenter and datacenter-adjacent bankruptcies before it gets better.
Some of these greedy bastards need to lose their pants on all of this.
You can run IQ3_XXS, IQ3_S and the IQ4_XS quants on this too. It works. It's fantastic. I'm getting better results than 27B now.
To add: https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RC... IQ3_XXS is within a point of the fully unquantized model and IQ3_S actually beats the unquantized model on many tasks! You lose absolutely nothing. It is quantization magic :)