Comment by nsagent

9 hours ago

See this recent paper: Quantization Degradation in Large Language Models: A Signal–Noise Perspective [1].

  We observe that such degradation varies substantially across these factors: 4-bit quantization usually preserves performance, 2-bit often causes broad degradation

This repo uses 2-bit quantization and removes some of the experts for its smallest fastest model. Make of that what you will.

[1]: https://arxiv.org/abs/2608.08188

I created a pruned experts model of the q2 quant, while it gave good performance on limited hardware there was severe quality degradation.