Comment by kennywinker
5 hours ago
Yeah, -10% accuracy (probably more like -20% in reality) sucks, but only if you could be running it at 100%.
That's the exciting part of this - before the best you could run on <24gb vram was qwen3.8-27b at q4 quantization. Now you can run a nerfed 125B parameter model on under $800 of hardware, and it beats a less-nerfed 27b model.
No comments yet
Contribute on Hacker News ↗