← Back to context

Comment by esafak

9 hours ago

Has anyone calculated the effective intelligence of these quantized models?

I think publishing benchmarks with quantized models should become standard practice.

See this recent paper: Quantization Degradation in Large Language Models: A Signal–Noise Perspective [1].

  We observe that such degradation varies substantially across these factors: 4-bit quantization usually preserves performance, 2-bit often causes broad degradation

This repo uses 2-bit quantization and removes some of the experts for its smallest fastest model. Make of that what you will.

[1]: https://arxiv.org/abs/2608.08188

  • I created a pruned experts model of the q2 quant, while it gave good performance on limited hardware there was severe quality degradation.

There's some info in the README, including:

> Coder: a coding version with half of the experts removed. It reaches 91% of the full model's SWE-bench Verified score (measured by its authors) and fits 32 GB of RAM.

https://github.com/Niko1221/Strata#which-model-should-i-pick

  • I wonder how this Coder compares to Qwen 3.8 27B. Can it be really better since they are competitive for same memory requirements?

    • 125b at q2 is ~80gb

      27b at q4 is ~16gb

      So from a raw amount of data, qwen3.8-flash-next wins easily. But flash-next is an MoE model, so it only has 6b parameters active per token, vs 27b's dense 27b per token. So 27b@q4 uses ~16gb of weights per token, and flash-next uses about 4gb of weights (125/80 * 6).

      But those numbers don't really tell us anything useful, because there is an interplay between total model size and active parameters and intelligence that isn't obvious or simple.

      (sizes are based on the unsloth quants, not the coder variant, but the idea holds - this isn't calculatable with simple math, you gotta test them and see)