Comment by kamranjon
7 days ago
After using a highly capable 2-bit quant as my daily driver for months now, I get pretty excited about releases like this. After a few days for the kinks to be worked out, I’ll be excited to try it.
7 days ago
After using a highly capable 2-bit quant as my daily driver for months now, I get pretty excited about releases like this. After a few days for the kinks to be worked out, I’ll be excited to try it.
What model? And what hardware do you run it on?
I find these style of models are great, but fail hard, and fail randomly. I'd be hesitant to use it for a daily driver, but I'm using dual 3060s, so it's not like I'm quantizing a frontier model here.
How do you find the overall experience? And do you have any special sauce or recommendations for going this route?
I’m using DeepSeek V4 Flash on 128gb mbp - it’s a bit different using a 200b+ param model. It’s MoE so performance is acceptable. It will still malform a tool call every now and then, but the capabilities are so far ahead anything else that the majority of the time it works really well and solves really complex problems.
I've got dual 3060s as well. What's the best models you've found for this setup?
What have you been using?
DeepSeek V4 Flash with DwarfStar: https://github.com/antirez/ds4
The 2 bit quants are really good. I have a lot of memory so I can squeeze it all in at ~80gb.
But isn't it running basically 1 request at a time? This would make agentic coding difficult right? Compared to running as many sequential tests as you want via api?
1 reply →