Comment by nullc
2 days ago
I wouldn't find glimmer interesting except that it has much less memory usage per token of KV than Qwen. So I can get 24x concurrent glimmer on 2xRTXA6000 (with 128k context) where I can only get 6 Qwen 27b. This means I can get something like 4x the aggregate tokens/s out of glimmer.
For some usages that speedup more than makes up for it being inferior to Qwen intelligence wise.
No comments yet
Contribute on Hacker News ↗