Comment by Jeeetendra

9 hours ago

getting it to fit is impressive, but i'd want to compare the smaller quants on a real coding task before picking one. how much quality do you lose going from IQ3_S to Q2_0?

I have fairly limited HW, so i tried standard llama-cpp and qwen3.8-flash-next, unsloth quants.

Q1 was producing some garbage at times, generating wrong urls on webfetch, then convinced itself there was some url rewrite in the middle. With IQ2 it happened much less but still happened, and once it would all webfetches became like that. IQ3_XXS is the maximum I can run: I don't have problems anymore, though I have less available context window.

  • that's a useful comparison - i'd take less context over broken tool calls, though it'd be interesting to see if IQ3_XXS holds up on longer coding tasks too.