Comment by manmal
3 hours ago
My biggest learning after some experiments - a BF16 (unquantized) Qwen beats a Q8 of double its size for decisions. I guess that’s the reason Kev switched to 4B BF16, from the original 8B version. Isn’t it interesting that quantization seems to mess with decision accuracy?
That’s fascinating, but not that surprising to me. We act like quantisation is free “Q8 is basically lossless” is often said in the local LLM community, but it really isn’t. The trade offs are worth it, personally, and the damage to coding ability seems low: decision model approaches are stricter though
Super cool finding!
Interesting - I wonder if it's because coding doesn't use the specific token probabilities, while decision models do
I think it also helps that speed and latency aren't as vital for coding and we can afford to let the model think for longer.
Yes this is the reason. Transformations like quantisation preserve the rank of outcomes much better than probability mass.