Comment by Tuna-Fish

6 hours ago

You don't want to use a sparse model for a Taalas-like design. Something like a Qwen 3.8 27B makes much more sense.

I'm not smart enough to know why; I do know that 27B is greater for short/interactive on blackwell, but the intellgence leap of the MoE in Qwen3.8-Flash-Next is quite remarkable.

I'm pretty convinced the pathway to local models will be MoE, especially if they can find a way to keep tweasing out things like PLE into the slow bandwidth lanes.