Comment by adrian17

12 hours ago

> Ternary Bonsai 2 27B uses ternary {−1, 0, +1} weights with FP16 group-wise scaling, for 1.76 effective bits per weight

If I recall correctly, a recent post [1] has shown that Q2 quants (with like 2.6 bpw) of the same base Qwen model sit at the edge between "noticeably worse" and Q1's "useless". I took a quick glance at Bonsai's blog posts, and don't really see them comparing themselves to "typical" quants or explaining what's the special sauce that makes them better?

https://news.ycombinator.com/item?id=49611128

There's a table on the HF page that compares it against UD-Q4_K_XL and IQ2_XXS (you need to expand the dropdown): https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf#fu...

The table claims it performs on par with UD-Q4_K_XL except on OCR.

  • I wonder how well it performs in practice, because I can't help seriously doubting those benchmarks. That would put this 6GB model in Opus 4.6+ ballpark. Granted, that's mostly to Qwen 3.8's credit, but it's hard to believe that Qwen's already unbelievable capability density can still be compressed this much more.

I think the general idea is naive quantization falls apart below 4bpw but you can go lower with more sophisticated QAT-adjacent methods. Bonsai's quantization method is proprietary though.

  • That's correct. I think there can certainly be issues even with 4bpw with naive quantization (IE you'll notice far better results from a QAT 4bpw vs a naive 4bpw).

    One such method that I've been meaning to look into further is Tencent's AngelSlim QAT/PTQ approach. They did a Hy4 preview release thats an STQ_1_0 at 2.38 bpw:

    https://huggingface.co/AngelSlim/Hy4-preview-GGUF https://arxiv.org/abs/2602.21233

    Of course, it's still 213g of VRAM I'd need so it's somewhat out of the range of what I can run locally. In contrast, this new Bonsai is nice because the original was already exciting for making use of low VRAM devices. Could breath new life into some of the older GPUs that were previously close to top of the line just quite VRAM constrained by modern standards and still quite cost effective for now.

  • Could've been better if GGUF implemented QTIP format. GGUF representation is a major limitation for llama.cpp quantization performance

1.76 bpw number is kinda misleading if you compare it directly to IQ2/Q2. The encoding is ternary, but the quantization procedure is way more sophisticated than "round Qwen weights to {-1,0,+1}."

They rotate the weights into a quantization-friendly basis first, then ternarize with per-group scales and error compensation.