Comment by tornikeo
8 hours ago
I love how the mere mention of a "sovereign" in LLM's announcement is the declaration of defeat.
This thing is worse than a Qwen3.8 27B.
8 hours ago
I love how the mere mention of a "sovereign" in LLM's announcement is the declaration of defeat.
This thing is worse than a Qwen3.8 27B.
Hi there! I'm Tom. I worked on Kolibri at Aleph Alpha :)
I think it's not simple to directly compare a 27B dense model with an MoE model like ours. As we know, dense models need all params active for every token. Whereas, MoE models (especially sparse ones like Kolibri) fewer active parameters and correspondingly less compute per token.
Among the MoE models we compared against in our tech report and model card, though, Kolibri performs very well in our evaluation, including against models with 12B active parameters. It also best model in the group we tested within that range of active params.
So, I think it's fairer to see this as a trade-off. Kolibri needs less compute per token but more memory, while Qwen3.8 27B needs far less memory and more compute per token. In the report, both are actually on the quality-vs-serving-cost Pareto frontier among the models we evaluated, just at different points.
It’s an incentive problem. If “sovereign” becomes your claimed value proposition, you can claim success even if the models not competitive, so nobody is pushed sufficiently hard to actually make it good.
Sovereign works when talking about building a commodity supply or something, not in literally the world’s most competitive and fast moving field.
Those seeking sovereign capability would be better off aiming to be best at something, even something much narrower than an all round LLM. Or just fast following and making something that matches leading performance, which is close to what the Chinese labs do currently.
I can think of at least one other pretty good reason, which is in anticipation of regulatory capture. If "LLM used must be FOOBAR-certified" and coincidentally no Chinese models can get this certification, having such an alternative is a lot more valuable than just scoring highest in a set of benchmarks. Not to mention that these benchmarks aren't always accurate.
It doesn't help that Qwen3.8 27B is an excellent model.