← Back to context

Comment by manlymuppet

8 hours ago

Man, a lot of this discussion sounds like people cheering for the last kid crossing the finish line.

Surely we want competition and Europe involved in that, but at this point I have grown used to either American labs smashing the frontier remarkably fast, or Chinese labs getting way, way closer than you would expect them to.

Mistral’s progress, regrettably, feels much slower. This model doesn’t knock anybody’s socks off. The model is (and I hate to be this harsh) mediocre, and this mediocrity has also arrived months late.

This is a pretty grim prognosis for European AI.

> Man, a lot of this discussion sounds like people cheering for the last kid crossing the finish line.

Sometimes it's ok to cheer for the last kid crossing the finish line because they're actually running a totally different race, and winning might look completely different.

When I look at what Mistral does vs other organizations I'm impressed:

https://isaiprofitable.com/

They aren't profitable yet, but they're a lot closer than most and they're doing a hell of a lot with very little.

Pointless racing story:

I was in high school track with a really tough guy who was just not a runner. We went to a pretty messed up high school and if you screwed around in track practice sometimes the coach would make you run a crap race at the next meet, like steeplechase or hurdles. Well this guy and a few others screwed up and coach made them all run hurdles at a meet.

He hooked every single one and fell on his face. Every time he got back up and kept on running. By the time he hit the finish line his knees were bleeding halfway down to his ankles. We cheered like hell and he was smiling ear to ear.

Coach quit punishing us with races after that.

  • I read that site quite differently from you. You seem to be analyzing absolute differences but ROI is really about ratio of spending to revenue.

    It looks like Mistral is middle of the pack, behind Anthropic and ahead of OpenAI on that front. All of those labs are way "ahead" of the cloud providers, but those providers are building infrastructure, not just training models, so it's not apples to apples.

    • Yeah that's a good point. The numbers are smaller and I was looking at it absolutely.

      Speaking from an absolute perspective I do think they are doing more with their money than either Anthropic or OpenAI.

> This is a pretty grim prognosis for European AI.

I think it's an incomplete read. What's the point in competing for a sizeable percentage of your funding when the finish line is incrementally being moved each month? Better spend it on leapfrogs which they seem to have done.

Meanwhile Mistral have a natural ace in their pocket with respect to regulation in the form of CADA and the Cloud Sovereignty Framework. I can't think of another company that would qualify as SOV-3 under that regime

  • So we’re supposed to cheer on companies that make worse products and only exist due to regulatory capture now?

    Political polarization is turning the world insane.

I don't think it will matter in a year or so. We are clearly topping out on useful intelligence for an increasing amount of tasks, as demonstrated by more and more models reaching the "useful" barrier.

This barrier is not going to start moving dramatically. It will simply be mostly satisfied for most work we do. Mistral is going to get there, soonish, long before the economy takes an entirely different shape (in so far that even happens).

There will be super human intelligence tasks, tasks truly constrained by intelligence for quite a while. Those will be few and far between, relatively speaking. Mistral will have plenty of opportunity to capture the other stuff, with a fraction of the resources required that it took the frontier labs to get there first.

People were extremely dismissive of chinese models until recently. They went from 1 year behind frontier to 6 months behind frontier to 3 months behind frontier extremely fast.

  • To be clear I'm not trying to dismiss European AI. I am a proponent of it.

    But Chinese models have very much earned their place. The same cannot be said of Europe, so far.

  • Serious people haven't been very dimissive of Chinese models since at least DeepSeek-R1 in January 2025. Sadly, we in Europe are much behind.

  • Important to point out that these 'X months behind frontier' really refer to the public frontier, and not the actual frontier, which private companies are free to protect indefinitely. Perhaps open models are in actuality 18 months behind the actual frontier - how would any of us know?

Curious, why do you say the model is mediocre? I haven't tried it, so I can't pass any judgement... I've learned to distrust benchmark rankings. Are benchmarks and Artificial Analysis the yardstick you're using?

  • > Curious, why do you say the model is mediocre? I haven't tried it, so I can't pass any judgement.

    It's nowhere mediocre.

    It's toes-to-toes with GLM-5.3 which is one of the best Open Weight model available (With Kimi K3) for general reasoning.

    I just runned it on code reviews right now and it was able to catch some thread safety issue than DeepSeek-4.1 didn't. And DeepSeek-4.1 is by no means a bad model.

Doesn't matter, it's excellent, and it's European with European inference, which solves the pains of all my clients trying to build data lakes and processes on top of it.

Nobody in the real world cares about minor benchmark differences in money losing coding agents.

And nobody in the real world is giving Altman or Musk their data.

  • I wish you were true - but I still meet a lot of people handling sensitive data and using free or cheap version of ChatGPT or Claude with their customer data.

    I think we will see some horror stories come out with data leak in the next years.

Neat. Wait 3 months for the landscape to change entirely.

LLM development is jumpy. It’s hard to extrapolate very far ahead.

  • I agree.

    When Europe does surprise us, I will be the first to commend their progress. But until then, this is where we're at.

First few models will always be slow improving and worse. The way to improvement is working your way through a gajillion evals [1], finding bugs, gaps, and curating training data (this part involves human design as well as raw inference compute) to fix it. This is very time intensive and can't easily be "done once and then everyone has lesser work to do" since every model is different. Well, one way to accelerate it is to simply have more compute, which mostly openai and anthropic have[2].

This is mistrals first 1T-scale model and I expect the 4th or 5th generation to be close to the best for many purposes.

[1] These evals differ from the public ones like terminal-bench, are sometimes model-specific, need real, diverse usage to actually create, and are held secretly since quality of eval is the first driver behind the next step improvement of a model.

[2] It is not close. This model was trained on less than 4k GPUs, whereas astra used north of 100k GPUs.

  • Mistral is not that new a player though. How can we give them this much grace when other players like xAI have done more in even less time? I don't think coddling Mistral helps them.

    And to the point of scale and training cluster, so what? Not only do Chinese labs have smaller clusters with less empowered GPUs, compute is Mistral's responsibility. You can't take away from other labs just because they fulfill that responsibility better.

    • xAI has a lot of compute. Deepseek also has a lot, not as much though. But this is changing with their new 160k huawei ascend datacenter in inner mongolia.

      The lack of compute is not really attributable in that sense to mistral. First of all it needs general investor and government willingness, which is easier in a larger economy like the US or China.

      Second, you need widespread usage of your paid inference service for two reasons: one it pays off your compute cost, and two it speeds up the improvement process.

      The vast majority of deepseeks paid customers are within china itself (since openai and anthropic services are not reachable from china) which gives it a market. But for someone in france, there is no reason to use a structurally slower developing model from mistral compared to using one from openai...except when data guarantees are needed, hence the landing page focus on sovereignty. As far as the dual use aspect goes, a model like this is more than enough, so the government will be happy.

Those are comments from Europe. The US is waking up now and I expect them to be much harsher.

I really want them to win as that's our last horse in the AI race, but ~200 research-oriented devs out of 1800 employees? I believe they agree it's pretty doomed and have pivoted.