← Back to context Comment by christkv 9 hours ago Awesome 3.8 next runs great on my Framework Desktop so I'm loving more local models. 3 comments christkv Reply hypercube33 8 hours ago I have similar hardware - what specific version of the 3.8 Next model are you running and how many tokens/s are you seeing? 3.6 35B A3B gets about 68t/s for me so I've been sticking with that model for the mean time. throwa356262 8 hours ago The thing with 3.8 next is that it uses a variant of ngrams. Part of the network is replaced by a lookup table you can store on a fast ssd.In practice, you will be able to run models a bit bigger than 35B.https://unsloth.ai/docs/models/qwen3.8-next cyanydeez 6 hours ago https://github.com/peonist-ai/halogen-server is beating the pants off of anything I've tried.38GB of vram resident. more tk/s, more prefill.
hypercube33 8 hours ago I have similar hardware - what specific version of the 3.8 Next model are you running and how many tokens/s are you seeing? 3.6 35B A3B gets about 68t/s for me so I've been sticking with that model for the mean time. throwa356262 8 hours ago The thing with 3.8 next is that it uses a variant of ngrams. Part of the network is replaced by a lookup table you can store on a fast ssd.In practice, you will be able to run models a bit bigger than 35B.https://unsloth.ai/docs/models/qwen3.8-next cyanydeez 6 hours ago https://github.com/peonist-ai/halogen-server is beating the pants off of anything I've tried.38GB of vram resident. more tk/s, more prefill.
throwa356262 8 hours ago The thing with 3.8 next is that it uses a variant of ngrams. Part of the network is replaced by a lookup table you can store on a fast ssd.In practice, you will be able to run models a bit bigger than 35B.https://unsloth.ai/docs/models/qwen3.8-next cyanydeez 6 hours ago https://github.com/peonist-ai/halogen-server is beating the pants off of anything I've tried.38GB of vram resident. more tk/s, more prefill.
cyanydeez 6 hours ago https://github.com/peonist-ai/halogen-server is beating the pants off of anything I've tried.38GB of vram resident. more tk/s, more prefill.
I have similar hardware - what specific version of the 3.8 Next model are you running and how many tokens/s are you seeing? 3.6 35B A3B gets about 68t/s for me so I've been sticking with that model for the mean time.
The thing with 3.8 next is that it uses a variant of ngrams. Part of the network is replaced by a lookup table you can store on a fast ssd.
In practice, you will be able to run models a bit bigger than 35B.
https://unsloth.ai/docs/models/qwen3.8-next
https://github.com/peonist-ai/halogen-server is beating the pants off of anything I've tried.
38GB of vram resident. more tk/s, more prefill.