Comment by c16
2 days ago
Big thank you to the Qwen team. 3.6 A3B was shocking good, and now I'm hoping they release an 3.8 A3B model too.
Edit: Having used qwen3.8:27b-mlx on MBP M4 64GB, I get around ~45 tok/s. A3B would be great for smaller devices, but it's definitely usable. As I understand it it's a mixture of MLX and MTP.
Looks like their A3B model is on the way
[1] https://www.reddit.com/r/LocalLLaMA/comments/1voxppd/qwen_38...
Thats a huge tok/sec. Prompt prefill is the bottleneck