Comment by abraxas
12 hours ago
I'm not following the local mdoel scene too closely but this seems quite amazing. Is this able to be run on Apple silicon too?
12 hours ago
I'm not following the local mdoel scene too closely but this seems quite amazing. Is this able to be run on Apple silicon too?
Their first 27B bonsai was able to run on an iphone.
"Ternary Bonsai 2 27B reaches up to 143 tokens/second on NVIDIA GeForce RTX 5090 and 46.8 tokens/second on M5 Max. On an RTX 4090, Ternary Bonsai 2 27B consumes just 0.714 mWh/token, making it 40% more energy-efficient than an 8B model running in full-precision."
Their mention of the 5090 is bit odd, since on 32 GB GPUs, Q6 fits while having better quality. Very interesting model for 16 GB GPUs though!
Sometimes you want a decent model running in the background that doesn't take up all the VRAM.
1 reply →
They mention 5090 with regards to speed, Q6 will not have that speed?
And speed matters a lot for many use cases
2 replies →
it is like "my fridge is 2mkm (millikilometer) from my desk" m=0.001 h=3600 it should be just Ws or just J
What’s wrong with milliwatt hours?
1 reply →