Comment by liuliu
7 days ago
The problem, of course, is if you run the UD_Q2 variant (Unsloth) which does only post-training, the number is pretty close to 1-bit model here and the 5% drop in tool-call is significant than it suggests in real-life use cases.
You also need to pay close attention to BFCLv3 multi-turn result, that helps you to get a sense how frequently these quants will be in a doom loop.
I'm curious what kind of results one could get from combining the clever quantization PrismML is doing here with something like LiquidAI's antidoom:
https://github.com/Liquid4All/antidoom