← Back to context

Comment by nekitamo

2 hours ago

This tracks, in my experience the 27B is better at coding and instruction following. I'm shocked at how much of a difference the dense models vs MoE makes.

But it's a moot point, because for local inference on consumer hardware, the MoE is so much faster.

Are you sure the difference is from MoE and not that 3.6 is newer?

  • Qwen 3.5 27B also scores higher than 3.5 122BA10B. So even in the same generation the smaller dense model outperformed the larger MOE