← Back to context

Comment by donmb

7 hours ago

Mistral is not that bad as the comments here suggest. I am not using it as a frontier model but with simple RAG tasks and its doing great. Also OCR is pretty decent. It's a positive development that Europe is at least trying. Alternative would be: do nothing.

The problem is that they're in a weird position between US models and Chinese models. Not as performant as US models, not as cheap as Chinese models.

Especially as Chinese models are getting better Mistral is getting less and less relevant.

It pains me because I want them to succeed, but despite them denying it I believe they'll end up restrict their activity to (1) selling hosting for Chinese models (they're already hosting GLM) and (2) selling AI-related consultant service (they're also doing that already).

  • I think Europe was at risk of being left without even a foot in the door. Mistral os one of few such feet. Just having some compute and know how for running Chinese models is better than nothing. Low bar I know but still. Who knows what R&D they’re doing while keeping the business running.

    I’m not sure really what the winning strategy is in this weird arms race.

  • If there’s one thing the legacy, dying industrial companies of Europe love, it’s consulting.

    So they’ll probably make more money creating PowerPoints with ChatGPT than they will trying to compete with the US and China.

    Which would be the most European outcome ever.

When it comes to what really matters, they're far behind and probably will not catch up. It is my conviction that in this space, if you're not the best in the world, you're losing. Everyone is fundamentally selling the same thing, so if you don't have the most intelligent or cheapest model in the world, you're losing. Sure, Mistral has a "Made in Europe" edge, but any open-weight self-hosted chinese model might as well have been made by Von der Leyen herself.

Also, €3B is nothing in this market, especially when you're competing with more efficient competitors. €3B in Europe is probably the same as €10B in the USA and €20B or €30B in China.

  • > any open-weight self-hosted chinese model might as well have been made by Von der Leyen herself.

    This is simply wrong. Plenty of security-sensitive companies and agencies disallow the use of any Chinese model out of hand. Given the black box nature of LLMs, the fact that it's "open weight" is irrelevant to the trust calculus, and the fact that it's self hosted barely adds anything.

  • > It is my conviction that in this space, if you're not the best in the world, you're losing

    I disagree on that point. Models are getting good enough that you can switch them and barely notice. I'm switching between Opus, GPT Codex and GLM 5.3 for coding and I can barely tell the difference.

    I think they'll become more like telcos than anything, selling a commodity. It's even truer when any provider can host open weight models like GLM-5.3.

    Basically a world with dozens of Baseten, with AI labs having a hard time monetizing, just like editors of open source software.

    • I can tell the difference between the SOTA and GLM 5.3, and paying a few hundred for the better models is definitely worth it.

      Mistral does not offer a model that makes sense to use. I understand they now host the already outdated GLM 5.2, courtesy of China providing the weights. And Mistral offers this for 2x-3x the price of other providers.

      This is supposed to be a success story?

      4 replies →

  • I don't get it when people all claim that AGI is a winner takes all game. It is not (unless it is used as a weapon). When one company reaches AGI, there will be a dozen very close to AGI, given time. Also a winner is not going to drive everyone else out of business, it is the opposite, one winner will have people betting on the second and the third winners, the technology will also help other develops. Once you have a good enough model, everything will be incremental. I think the hardware capability will be the real burden, not the model itself. If AGI is as powerful as it sounds, maybe hardware won't be a problem any more.

    • It’s very possible that the people saying AGI is a winner takes all are indeed referring to warfare.

  • Europe is a different beast. They build and distribute rails, they don’t compete on frontier capabilities. When the civic benefits of AI become clear, EU is in a position to mandate their distribution. US develops capabilities that remain stuck in heterogeneous corporate silos without interop. Payments is a good point of comparisons between the two approaches.

    • A new ICE line between Antwerpen and Koln opened up the last week. The first train was half an hour late, the distance is under 250 km.

  • Well, then Samsung Electronics appears to be dumb for leading this investment round? They probably only do it because they're European...oh, wait...

    > so if you don't have the most intelligent or cheapest model in the world, you're losing

    So how does this match up to the fact that there is currently OpenAI and Anthropic, both raking in money? They can't both have the smartest model at the same time, can they? And all those inference companies selling API access to open weight models on OpenRouter, which are apparently also earning billions already? While the former are probably bound to have much higher cost for research and training than they are currently earning, which may be called "losing", the latter don't have that problem, they can simply price their API access such that the money earned covers their costs, no training and practically no research necessary. In your theory these companies shouldn't have a cent of earnings.

    • > So how does this match up to the fact that there is currently OpenAI and Anthropic, both raking in money? They can't both have the smartest model at the same time, can they?

      They are simultaneously first: the two leapfrog each other with regularity, and are meaningfully ahead of the competition.

    • OpenAI and Anthropic are so close on benchmarks and release date that they are essentially ex-aequo at this point. The market is just hedging their bets.

      There is no close second, because the AA Index points are expentially harder to get as you get closer to the first.

    • The value of these companies is not defined entirely by the state of their current models. A much more important signal is their chance of having the best model in the future. Just like with anything having to do with investment, this is the castle in the sky. Also, shame on me for simplifying things so much - but I still believe in my original sentiment.

      1 reply →

    • Samsung is now essentially tied for being the world's most profitable company with Nvidia ($62b operating profit last quarter, vs $63b for Nvidia). They're drowning in cash. They're doing the same thing Nvidia has been doing: distributing the money to their business partners, to try to drive business growth faster (and or to keep it all propped up).

> Europe is at least trying

Which is interesting since who are the investors and what exactly are their roles? Samsung - European? BlackRock - European? Salesforce Ventures - European? Etc.

So yes it might well be

> [...] the largest equity fundraising round ever completed by a European technology company, three years after the company's launch.

but the money isn't European.

Mistral is odd. They have made mostly flops, boring models (Ministral 3, Mistral Small 4, Small 3, Small 3.1) together with a classic masterpiece Mistral Nemo and very good Mistral Large 2407, Mistral Small 22b, Mistral Small 3.2.

That's kinda damning with faint praise, but I agree, I am glad to see something. I've been disappointed in their coding ability -- it's where I'm most focused, I've built and we are selling a (specialised) coding agent -- and hopefully investment will give them the ability to achieve more in their research and model development.

They have been focusing largely on government and business not consumer, which is fine. Perhaps coding is not something they want to achieve, but they do provide Codestral. It's a signal it's a market of interest to them.

"Do nothing, win." is already China's strategy after all. Though they are far from doing nothing in the LLM space.

  • China is actively tripping the US up by pushing the open-weight strategy, which is kind of smart. They realized that they obviously will not be able to make Western companies trust Chinese companies enough to just use proprietary models by sending all their data to services hosted by Chinese companies. The US is still able to attract this trust, although it's eroding quickly in spite of recent political events and the current trust is more or less a function of old habits that need some time to change. But China correctly found that - due to there being no old habits to piggyback on - they wouldn't ever be able to compete in this way, even if they had proprietary foundation models superior to their US counterparts, so they decided to throw sticks into the spokes of the US frontier labs by releasing top open-weight models worth billions of dollars in training cost and thereby devaluing the huge proprietary investments of the US labs.

    That is absolutely not "do nothing".

    • Their ‘open weight strategy’ was invented, to Chinas surprise, by the Western press after the DeepSeek event. Of course all the actual systems like Baidu, Bytedance are closed weight and basically all anyone uses. Meanwhile the advanced system are crossing security frontiers: soon only inferior models will be ‘open’ (as with OpenAI, even) the rest closed. — Or where open weighted, they will be impossible to run without a private data center, come with absurd ‘security scrutiny’ licenses like GLM is starting — or licensing requiring a cut, as with Kimi, which is basically a scheme to get western companies to do buildout for them

  • Europe is not China, I repeat we are not China, they build stuff. A guy from Austria ended up on Chinese national TV due to the bizarre way he had to obtain an AC unit this summer. Let's stop the larp.

Europe could tell ASML to put kill switches in GPUs so Europe has leverage to "safeguard" AI deployments in other countries.

  • Ehm, sorry, but no, this is not how lithography works. You cannot "hide" functionality in circuitry your machines are producing if your machines' job is to shine light through a mask.

    You'd have to be the producer of the machine producing the mask. Or, even better, the software that produces the plan according to which a machine produces a mask.

    • I didn't say to hide it. My suggestion is to create international AI safety rules which companies must be in compliance with if they want to buy and service ASML machines.