Comment by Tuna-Fish
6 hours ago
You don't want to use a sparse model for a Taalas-like design. Something like a Qwen 3.8 27B makes much more sense.
6 hours ago
You don't want to use a sparse model for a Taalas-like design. Something like a Qwen 3.8 27B makes much more sense.
I'm not smart enough to know why; I do know that 27B is greater for short/interactive on blackwell, but the intellgence leap of the MoE in Qwen3.8-Flash-Next is quite remarkable.
I'm pretty convinced the pathway to local models will be MoE, especially if they can find a way to keep tweasing out things like PLE into the slow bandwidth lanes.