← Back to context

Comment by andy99

3 hours ago

What’s the point of having “sovereign” weights that are worse than publicly available ones? Wouldn’t Europe be better off just keeping up-to-date on the Chinese releases? In the event of some schism requiring sovereign capability, or even if the Chinese pulled ahead and stopped releasing the weights, why would Europe be better off because of Mistral? (Or any country’s inferior sovereign effort make them better off?)

I think I understand the incentives that cause this to exist (it would be politically worse to say we’re just going to use Chinese models) but they are misguided. If sovereigns want to have valuable models, they should insist on world class, relevant ones like the Chinese have. Instead they embrace mediocrity in the name of sovereignty.

> In the event of some schism …

In the views of most Europeans, that schism already happened.

Europe was perfectly happy to rely on US software and services for decades. None of the large US tech companies would be nearly as profitable if they hadn’t had a whole continent of wealthy customers, and no competition.

I don’t think Americans are realizing yet how much has changed for us the past two years.

2 things here, one is related to benchmaxxing, another to being good enough

1. Mistral isn't benchmaxxing. That doesn't mean they're better, but it does mean the benchmark gap not a good reflection of the actual gap

2. I think the "world class or nothing" framing mixes general capability with system capability. Most deployments don't need AGI In RL you need a model that's reliably good at one or two things, thats it. Example: Case of a hospital flooded in emails. You make a system that decides which patient emails needs a human and drafts replies for the rest. If a sovereign model is good enough at that, and you can run it on a hospital's own servers under EU jurisdiction, the frontier gap part has zero importance

Who cares about "beats DeepSeek / GPT11 / Claude Fairytale 8.9"

  • Agree benchmarks don’t tell the whole story. A better and easier evaluation of capability is whether anyone is using it for anything.

    Does Mistral have material market share for any application, including anything that would fall under item 2 above?

The training data and knowledge is the edge, you need to build that up and maintain it. And of course mine everything you can from the American and Chinese models, like they mined everything from the internet / films / music / games etc.

Not sure if you are European, but in EU it's a bit taboo to even talk about this in this manner. We like to spend a lot of money to make sure we finish last.

We know it's possible to put backdoors into LLMs, we don't have reliable ways to detect them without direct support from whoever inserted it.

Europe is less-worse-off with open weights than with… I guess it's weights-as-a-service? WaaS? The thing Anthropic and OpenAI do.

But that's not enough. As recently demonstrated, being just a few months behind with the power differential between defending with an open weight model while being attacked by a leading model, means losing absolutely.

I do not know if this holds going forward or not. It's not inconceivable that we're just about to get models that make unhackable code, using all the things software developers keep saying you need to do if you really care about security.

But anyone concerned about sovereignty can't bet the farm on this possibility. For the moment, it looks like it's a national security matter to ensure at least core state functionality (including core private sector logistics) gets the absolute best attention money can buy, and that the absolute best money can buy ("can buy" does a lot of heavy lifting here) is currently LLMs, and it's important those LLMs aren't going to get cut off by arbitrary whim like Mythos was, and it's important that those LLMs don't have backdoors like we can't rule out anyone else's from having.

This is true even if Europe was only defending from Russian cyberwarfare and didn't need to plan for the president of the country in which Mythos was developed, attempting to annex two NATO states.

Any model, even an open weight one, is fundamentally an encoding of a way of viewing the world.

What kind of "alignment" are AI labs optimizing for? Ideological alignment is the full term, self-censored into something more technological-sounding.

Every model has people behind it rating what it should and shouldn't say. Every time you ask a model and trust its answer, you become ever-so-slightly ideologically indoctrinated.

I don't want my model to reflect the views of American oligarchs or Chinese cadres. I want European values of enlightenment and humanitarianism to be the default and that's why the sovereign part is important.

  • Is there a specific concern you have, and what kind of performance penalty is it worth to you on say coding tasks?

    Conceptually, sure I understand, but in practice it currently seems like it amounts to just using a worse model without getting anything in return. And if some hypothetical alignment to European values is important, it seems like putting the necessary effort into building a model that’s actually competitive but has this alignment is the solution, rather than accepting an inferior one.