← Back to context

Comment by Balinares

1 day ago

I've found it tends toward long thinking loops even for simple tasks (and any quantization seems to increase their length), but those do exit eventually, unlike with Qwen 3.6.

I use the Unsloth UD_Q2_K_XL GGUF with default parameters, along with that custom template linked elsewhere in the thread, and no K/V cache quantization.

For smaller models, you'll probably find anything below Q4 will need handholding. Check Unsloth's graphs at the different quantisations VS error rates and you'll see why.