← Back to context

Comment by Tepix

5 hours ago

All headlines about LLM performance MUST have the quantization also mentioned in the headline.

You know, so you're not wasting your time like in this post.

Quants vary by model. DS4 is very credible at a 2-bit quant. Not sure about Qwen 3.8 Flash Next; I run it at a 4-bit quant and it's too slow, so I'm trying out DwarfStar today to see if that improves things.