Comment by spijdar
9 hours ago
The absence of any comparison to Qwen3.8 Flash, another MoE model with a small-ish (6B) number of active parameters, is pretty striking. Instead, it's compared with Qwen3-Next 80B-A3B, a model released almost a full year ago.
I get that doesn't invalidate the real "point" of the model, but...
That is what stood out to me as well. Qwen 3.8 Flash can run on very limited hardware as well and it is at least as good as Sonnet 5 in benchmarks like DeepSWE. With its n-gram design you can get flash next running on very limited GPU resources, as little as 16GB of VRAM.
People follow the latest frontier lab models with great attention and migrate to the next big model on their subscriptions. Meanwhile these local models have quietly gotten REALLY good. It is not even an exaggeration. It has happened in the last couple of months.
"Local model you can run at 40 t/s on a gaming machine that is better than Opus 4.6" is way less exciting than "OpenAI IS DOING CRIME!!! OpenAI SOLVED NAVIER STOKES. DARIO SAYS GLM 5.3 BAD! SLOW DOWN THE FRONTIER!".
(edit: also... totally ignore that 27B dense column over there where Qwen 3.8 27B beats Kolibri on nearly every single benchmark. Why would I choose to run this model?)
I think the simple reality is that if your goal is producing quality work output, you want the smartest possible model available.
What is the upper bound on the value of more intelligence applied to your problem domain?
The point of a model like this is aimed at providing local inference. You want the smartest model available, as long as it meets all of your other criteria. That isn't always maximizing on the absolute smartest model. There are many reasons to avoid a frontier model for now for a variety of reasons. All of this is even only relevant in the last 6 months anyhow. So it isn't like this is even some long term trade off or position I am proposing.
Can't you infer the comparisons you would like from baseline results provided? There are better results in the sibling post on HN: https://aleph-alpha.com/en/blog/kolibri-has-landed-a-soverei...
Maybe?
Qwen3.8 27B scored notably higher in most of the provided benchmarks, including the German-specific ones. The only "downside" is that inference is much more costly and slow, since it's a dense model.
Qwen3.8 Flash-Next appears to usually "benchmark higher" than 27B, while remaining fast.
I'm sure I could dig up the equivalent benchmarks for Flash and do the comparison myself, but as far as inference goes, it's messy. Consider that Qwen3.5 35B-A3B scores higher than Qwen3.6 on some of the German-specific benchmarks.
So it seems superficially plausible that Qwen3.8 Flash-Next might not be "27B but faster" in the ways that are important for this model. Or it could just "be superior" in all ways.
Either way, I don't think an LLM has to be "the best" at anything to be worthwhile, necessarily. And I kind of distrust benchmarks on top of that, so...