Comment by acd
17 hours ago
I have written a llm compression quantization library called glq because of the high ram prices. Glq uses qtip Trellis quantization which are efficient at low bpw 2-4 bits. As a gamer and ai developer I want to squeeze more out of the same hardware. Glq runs with vLLM.
Open source https://github.com/cnygaard/glq
Pypi glq https://pypi.org/project/glq/
No comments yet
Contribute on Hacker News ↗