Comment by stymaar
7 hours ago
I don't share your opinion:
IMHO, xhigh makes sense if you want a slower Opus4.6 at home. It is able to complete tasks autonomously in a way that I've never seen another local model do.
But yes, for simpler tasks or more hands coding sessions, it's simply not the best model out there as its verbosity makes unbearably slow.
Not just its verbosity; also its architecture.
Muse Glimmer has some interesting and it seems reasonably daring trade-offs in its architecture (that I wish I understood better) that seem to favour longer agentic “dialogue”, and it is just much more nimble all round, even though it’s a larger model.
Don’t get me wrong, I have spent time speccing out a box that I could use to run Qwen 3.8 27B better, and I am very glad it exists, as I am with the Gemma series. We have really an embarrassment of riches at the 32GB VRAM level already.
I just think maybe Meta have the more appropriate strategy (can’t believe I am saying this) for desktop AI.