Comment by xyzsparetimexyz
7 days ago
That's awesome. What's the largest model that could fit onto a single 16gb gpu at 1.125 effects bits per weight?
7 days ago
That's awesome. What's the largest model that could fit onto a single 16gb gpu at 1.125 effects bits per weight?
Doing some naive math, the F16 filesize is ~53.8gb, the 1-bit version is ~3.8gb, about 7% of the original size. The F16 size is roughly 2x param count, so that gives a rough ballpark of ~110B.
Which would be very interesting to test, as larger models (such as Deepseek V4 Flash or Qwen 397B) seem to compress better. Their Q2 quants are usable as is, even without the ternary compression.
Yep, that’s the question. I asked just that when Bonsai’s first models got released. Super interesting if we can push the parameter count over 100B with 1.125 bit quantization and still keep pretty good performance versus 16-bit 100B models. That’s a definite sweet spot.