Comment by anyfoo
17 hours ago
I have a very interesting self-made coding benchmark, very intricate and technical, but 100% a real world problem I had to solve. I’m not going to further elaborate, since I don’t want future models to train on the solution.
To my own surprise, Q6_K_XL (from unsloth) comes up with a solution, anything Q5 doesn’t. To further surprise me, so far only the XL Q6 variant managed to solve it.
The problem, at least as stated, seems to be right on the edge of what the Q6 quantization can do.
Unfortunately even a successful run is rather long, so I don’t have a whole lot of data.
But the whole thing sure made me doubt the common idea that you wouldn’t perceive a difference until crossing past 4 bits quantization.
I thought the wisdom is more so don’t bother going below 4bit and you won’t see a difference above 8bit.
Depends on the actual audience, I guess. My stated “wisdom” comes in part from /r/LocalLLaMa, and my impression is that the tasks that users there give their models to try them out lean towards rather simplistic, on the reasoning side.
But there I literally did read “you don’t need anything better than 4 bpw” a bunch of times.
> my impression is that the tasks that users there give their models to try them out lean towards rather simplistic, on the reasoning side.
That sub is games, porn, and complaints about not having money. A waste of time.