I wonder if the other models were worse than usual for that particular video because they "dislike" making an advertisement for another company's model. A test with a generic video might be more meaningful.
These are pretty much what a human would draw. Sun rises from the east. A cloud makes the background “sky”. Three lines is the minimum to interpret as movement. Two feathers is standard on every cartoon and illustration.
This comes up every thread. I think we've all noticed how similar they are becoming.
I would guess that they've definitely been trained on previous results, as they obviously share way too many traits at this point to be totally random. That said, I don't think we're seeing any signs of pelicanmaxxing yet from the providers, so it's still a useful (or at least fun) benchmark.
Once all the models produce pristine, elaborate pelicans riding perfectly drawn bicycles, then it'll be time to move on to pigs driving a racecar or something.
I have enough projects, I think some benchmark/dashboard showing kinship based on these kind of queries could be very interesting to watch and insightful when new models come out.
I think they're still visually pretty different. The most common shared details are:
- Pelican cycling to the right - that's been discussed at length, images of bicycles online always show that side of the bike because that's where the chain is.
- Bicycle is usually red. No idea! Red ones go faster?
Not bad! I like how it got the motion lines on the correct side. IIRC, many of the other ones you've posted have the motion lines on both sides of the pelican
The bicycle looks significantly better in high. And feet and hands are actually where they should be. The road looks worse, though. No flowers either. And in neither is the pelican sitting on the saddle, but I can understand it's hard for a pelican to ride a bicycle properly.
Now what would have been cool is if Mistral on high reasoning had realised that pelicans are the wrong proportion to ride a bicycle, and had designed a bicycle more suited to pelicans. Let me know if any model ever manages that.
High one is actually much better. The feet connect to the pedals, the wheels don't have a hub cap, although it looks like the pelican is wearing the seat, it's in a relatively proper position etc.
Both are riding on the left side of the path for some reason.
I have a question, and perhaps some of the AI/ML infrastructure experts here could answer: Mistral says, "ML4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own datacenters in Europe."
If a 1T params model trained on ~4k NVIDIA GB GPUs could almost match the performance of Kimi's K3 (which is on par with top closed source models of OpenAI/Anthropic) while beating/exceeding other leading SOTA models from top Chinese labs, what are we (in the US) even building these super massive data centers for? Just to churn through more backpropagation reps more quickly?
SpaceXAI's Colossus supercluster in Memphis and Colossus 2 in Memphis/Mississippi (Southaven) are supposed to run into hundreds of thousands to a million GPUs. MSFT's Fairwater GPUs are supposed to have hundreds of thousands as well. So, 3800 GB GPUs are an absolute drop in the bucket. I don't understand the strategy of hyperscalers here, especially with edge inference hardware only getting better from here on (Apple, and all).
Distillation explains some of the advances, but doesn't that mean hyperscalers have a ton of deadweight wrt GPUs sitting on their balance sheets? Will all these GPUs be used for inference once a SOTA model's training checkpoint/batch is done? It's bonkers to me.
> Will all these GPUs be used for inference once a SOTA model's training checkpoint/batch is done?
I have no real data to back this up, but that has always been my assumption.
Claude says that K3 can be assumed to have required 10-100M GPU hours. If you have 100k GPUs that would mean like 6 weeks of training. 100k GPU's can serve 3-30 trillion tokens of K3 per day. Google apparently serves ≈100 trillion per day [0].
The big labs probably want to have capacity to fairly quickly train / post train different SOTA models continuously + being able to serve peak inference demand in valuable markets (US daytime?).
It's unclear how big a role distillation plays, but it may be a big one. There's also a law of diminishing returns. To get a meaningful increase in model quality you seemingly need an exponentially larger model. And most people don't think K3 is actually on par with top closed models.
Google says 1 million blackwell gpus are being delivered monthly I'm really curious if it's companies just hoarding chips/memory/servers awaiting to be deployed in data centers not ready yet for months or years, or everything built is actually deployed upon delivery. Plus google and amazon have there own chips in the mix.
Today's data centers are being built for yesterday's inference need. There's a persistent cult belief that ai hasnt found a niche or that companies havent proven utility or use cases or whatever. the demand for ai (internal to hyperscaler, and external for everyone else) simply dwarfs what is available.
Impressive vision benchmarking. If the vision model is truly as good as astra, that would make it best in the world.
Also strong on cyber benchmarks (better than all chinese models), so this is a good defender model.
Lots of people shitting of Mistral for no reason imo. These are pretty good numbers across the board. Definitely good enough to use as a daily driver over other llms, if you have moral qualms with the others. For certain use cases, like cyber security, this may be the go to model.
I like to make fun of europe, but there's lots for mistral to be proud about in this release imo.
> Impressive vision benchmarking. If the vision model is truly as good as astra, that would make it best in the world.
Sure, it leads in one unpopular benchmark with internal numbers.
But look at the third benchmark. It only has Sol 5.6 from OpenAI and it shows a 12.8. I don't know this benchmark but AA's GDP.pdf listing has Sol 5.6 Max at a 27, even non-reasoning beats the 12.8. It's extremely weird to cherry pick Sol 5.6 and then also lie about the published third party benchmark score. I'd love AI companies to stop lying about this stuff.
I'm still embedded in the OpenAI ecosystem, but man, when I try out Mistral, it is snappy! Like basically instant responses that seem faster and better in quality than ChatGPT on instant mode. I'm impressed.
The paradox between supporting consumer rights with actions like universal usb-c adoption, but also complete elimination of any privacy rights at all. Like the surveillance state is insane. No e2ee chats, backdoors in everything.
That they think their politicians are better than the United States but have populist demagogues, right wingers winning, corruption scandals, freedom of speech restrictions, and unbelievably bad spending problems.
Still waiting for France to send troops to Ukraine or ask the United States to get rid of their European military bases.
Europe has spent the last twenty years in stagnation - generating about half as much wealth and technology as you'd expect for its size and advanced economy. Simultaneously, Europeans are notoriously arrogant. It's a mockable combination.
Edit: unfortunately, it's a question that voting does not really permit you to answer on this website.
> Curious, what do you like to make fun of Europe about?
Well I'm european and... That Switzerland (Europe but not EU) has more companies in the Top 70 by market cap than the entire EU (Switzerland has two, the EU only has ASML) is kinda something that warrants making fun of.
That the biggest European software company is SAP, in 71th position is both sad and tragic: it shows how lame and irrelevant Europe is when it comes to software.
So Europe is nowhere in software and friggin nowhere in hardware: sure it's got ASML but ASML now has officially... Zero customer in Europe. Zero is not much.
Then Japan is at least trying to come back into the game with nano imprint litography. Europe is betting it all on AMSL (which anyway is majoritarily US-owned).
So software: nothing. Hardware: nothing besides ASML.
Overall the EU has six companies in the Top 100 by market cap and they're all, besides ASML, near the bottom of the Top 100.
We could also maybe make a bit fun of how the EU destroyed it's car industry (the main industry in Germany, which is the biggest economy of the union) by handing it all to chinese EVs?
Or what about the US warning the EU, years ago, to not become entirely dependent on Russia for energy? And EU not listening and then seeing its energy price skyrocket when the proverbial shit hit the fan? (Russia attacking Ukraine)
And we could, also, at least make a bit of fun of entire streets in cities like Paris and Brussels that used to have luxury shops and fancy restaurants that are all turned into places selling cheap kebabs? What a great success: I'm sure this one makes the komrades happy. It projects an image of grandeur and success: kebabs.
Or the constant attacks on free speech in the EU. Or the surveillance apparatus that's being put into place.
And let's not forget: there were promises made to Russia to never grow the EU to the east. Then the EU started exciting Russia by saying they'd incorporate Ukraine into the EU: I'm not against that but doing that did trigger a war. And now suddenly the EU is waking up and feeling all warmongering, wanting to dedicate a big percentage of its spending to weapons and tanks and missiles.
The warmongering tiny pet that the EU is is kinda laughable too.
At this point it's more like I don't know what is there left to not make fun of about my EU.
If the UK is in europe, you should make fun of the following: they don't have a first amendment. So to me, it's an authoritarian state preaching freedom.
Even if it's not the best model, it can be really important step in UE sovereignty. Trained in EU, inference in EU. I guess it will matter for some companies. Hope Mistral won't disappear for the next half year.
the greatest threat to Mistral is that there isn't a deep and large enough capital market in the EU to absorb the valuation step ups needed by the handful of EU sovereign growth investors to justify their existence.
you could say, that's perpetually a tomorrow problem, so long as they never go public, but that should illuminate for you: if everyone "knows this," it's inevitable that this so-called EU company, that didn't invent any of the AI, the hardware nor the training data, where their product is essentially more like a VPN provider than a frontier technology company, just lists and capitalizes in the US anyway.
By your measure anyone unable to build EUV lithography machines without ASML help is doomed. EU could easily corner the market and shut them all off. No more NVIDIA.
The last 3% of the stack? They’re essentially borrowing Chinese open source distillation/innovation off the US frontier running on US/Taiwanese designed chips and calling it EU made. This is better than nothing of course.
But is sovereignty really an end in itself? So the EU becomes IT independent…cool, then what? I mean the Yugo was sovereign, it didn’t do much good when the society itself failed to produce prosperity and collapsed.
It seems to me the EU is expending enormous effort on the appearance of “sovereignty” over what is ultimately…the last-mile consumer toilet paper purchasing conversation data…of an aging, increasingly irrelevant population on the political/economic stage.
Meanwhile domestically the entire economic model is failing and the “union” is getting shakey as its 2 biggest members turn more nationalist.
Maybe this “oh you cant compete but here’s a trophy for sovereignty” attitude isnt helping. Less clapping along with the EU bureaucracy’s latest make work projects like transitioning to a new Word processor. If we want the EU to succeed, more tough love is needed imo.
it's industrial policy, which is of course the minimum for being able to even think about having independent thoughs when it comes to international relations.
hard to stand up for EU (or even national) values if some dude in Washington has the keys for most of your military
having at least some in-house expertise is the first step.
A strong competitor in cybersecurity as an alternative to GLM-5.3 (Mistral reports 82% on CyberGym-E2E). Visual grounding is also impressive (42% on Dense 200 vs. 41% for GPT-6 Astra).
Otherwise, behind on the broader Pareto frontier, but not by much (Vals Index: 48.05% vs. GLM-5.3’s 53.51%; $13.78 vs. $7.25 per test). Many companies will prefer it over Chinease models.
That was about my conclusion as well: Slightly less than GLM 5.3 performance but made in Europe. So, maybe it answers Tiananmen Square questions correctly, and in French. All in all, a reasonable model, but not frontier.
> So, maybe it answers Tiananmen Square questions correctly
What would you consider a "correct" answer? I just asked deepseek-v4.1-flash asking it what happened (without mentioning the word "protest"; here are some excerpts of what it said:
> In April 1989, students in Beijing began demonstrations after the death of Hu Yaobang, a former Communist Party general secretary. The protests grew. [..] Estimates from other sources range from hundreds to several thousand deaths. [..] The Chinese government describes the events as a counter-revolutionary riot and says the military action was necessary to restore stability. It restricts public discussion of the events inside China. Many other governments, human rights organizations, and observers describe the events as a violent suppression of peaceful protests.
So, let's see... it calls it a "protest", mentions the number of deaths, and even mentions the censorship of the topic by the CCP.
"..answers Tiananmen square questions correctly.."- but lies about Ukraine, EU, and about you, americans..
I live in EU, use for my personal needs chinese models only, and don't plan to move to any of the ones allied with the Pentagon or its european counterparts.
Its quite interesting to see that at least the early days of AI so far have not been a winner-take-all runaway acceleration game where catchup is impossible.
I certainly wouldnt have predicted that 10 years ago.
Very glad to see Mistral still in the game even after some big stumbles with Large 3. I deeply hope that this model is 'good enough' that it becomes the European go-to, giving them the resources to keep the pace up.
I think it might have accelerated things but on a much more basic level, there seems to be no real moat in synthesizing the world’s knowledge into LLMs.
My spend on DeepSeek is not much and I regularly top up my balance every month as my support for all the good work DeepSeek is doing for the open science.
Architectural/algorithmic tweaks do advance the efficiency frontier nicely. But raw intelligence mostly comes from data (not just its sheer quantity, but also how it's curated & cleansed) and the scaling law. The know-how about data curation doesn't seem to get published much, even among the open-weight labs, though.
I'm not sure I like this framing - so much of AI research has been academic, in the open, building on others people's work. Much less comp sci generally, math & philosophy, etc. The idea that rich companies can just build stuff in secret because they have resources is a fantasy.
Also, the field moves fast, but slower than people do. Researchers and engineers switch companies every year or two, and the know-how walks out the door with them.
You've got to think the comments about "good enough" are people who have not yet tried Opus 5.5. I haven't been this struck by a step change since 4.5/4.6. It's a bigger jump even than when Fable first arrived.
As for Mistral - I got really excited when they said Large 4 was focusing on being #1 in cybersecurity, because that's somewhere that they genuinely could edge out Anthropic & OpenAI. Have it actually solve problems, instead of Anthropic flagging "you tried to find a null pointer exception bug in your own code, we're now reporting you to the US government". But on the Mistral benchmarks I'm seeing, this looks very disappointing, but at least they haven't entirely given up. I genuinely thought Mistral had given up on new general models. They need to learn the bitter lesson all over again.
I agree that Opus and GPT are almsot surely better, but so many real users are nervous enough about giving Anthropic and OpenAI access to all of their internal information that they may be willing to stomach worse models if it gives them more security.
The real question is if this model is good enough that it can still accelerate work, and not be a hindrance to real work like older Mistral models often were.
Please, give it another 6 months and they catch up. The American labs are currently trying everything they can to block others instead of advancing their models, trying to build an artificial moat. The American models are not that great, they are good, and they have a lot of agentic workflows in the back, but its basically a hardware limitation at this point. Once the HW makers catch up, and we can move away from the Nvidia monopoly, things will speed up quite a lot IMO.
It's an improvement, but game changer might be a bit of a stretch. If I lost access to Anthropic or OpenAI models tomorrow, I would be annoyed, but would reach for a slightly inferior model. Last year I wouldnt be able to say the same, and rhe challenge is that the moat is drying up fast. Whether its general improvements in model training by other competitors, or straight up distillation of SOTA models, the moat is shrinking and the available capital and spend for American model providers is going to dry up quickly as competing good enough models are adopted by more consumers.
It's especially the case as more non-Americans look to self hosted models and domestic cloud inference providers using open models that the US providers who are still leading the charge need to drastically drop their prices and find a path to profitability in order to maintain their lead and retain the advantage they had as AI turns into a commodity (which is happening faster than I think even the frontier labs initially predicted).
I hear this literally every other week about whatever the newest FoTM model is.
Unless you can provide concrete examples of things you can do with them that you simply couldn't do with last week's model, it's absolutely meaningless.
Yeah, I agree with this. I think the "the models are good enough" narrative is a myth. I've heard it so many times over the last year, but the model number keeps changing...
There is no ceiling on what you can accomplish with more intelligence, so there will always be a market for the best models, and that market is likely to just keep growing. If Opus 13.5 can one-shot a profitable company or discover a new disease treatment or whatever you can think of that a swarm of relentless super-geniuses could accomplish, companies (and governments) will throw money at it.
I also think there will always be a market for many sub-frontier models that will continue to grow rapidly as well, because "good enough" is definitely a thing for a given task.
It's shaping up to be much more like a game of 'chicken' where each company tries to raise more cash without going bust...
Ultimately the game of musical chairs is going to have to stop. In the US it looks like they are trying to get a government sanctioned truce in the form of regulation. That's what 'Pacing the frontier' means...
Even "runaway acceleration" isn't instantaneous. People imagine the singularity as something that happens almost instantaneously. But obviously it happens over time, and that time might be decades. It might still end up looking like a vertical line on a long-term graph.
If the singularity is defined as an AI sufficiently intelligent to improve itself independently, that AI is still limited by the resources required to do this improvement.
This appears to be roughly as good as Sol 6.1 (which is quite good), considerably faster in terms of wall clock for complete tasks, and considerably cheaper (where Sol 6.1 is already good value - just really slow).
I can confirm that you're reading this incorrectly. There's a reason behind them only comparing it to open-source models released months ago. Here's a good aggregator: https://artificialanalysis.ai/#intelligence
It will become winner take all when AI companies manage to really get value from user logs.
Right now they don't even get good feedback from local sessions - I can see it make the same mistake two days running, and then months later when a new model comes out, presumably trained on my data, it still makes the same mistake.
Hear! Hear! I really want European models / AI labs to succeed.
I trust them and their populations to provide a more societal-friendly version of AI, putting pressure on the US tech oligarchy, while also providing democracy-friendly open models that I don't trust to happen with the Chinese labs.
Truly, this is what the Lord's prophets have revealed to us! (Eliezer 11:52) Keep strong in your P(singularity), for when the Kingdom arrives, He shall judge us in His righteous glory, whether to eternal annihilation, or rebirth and life in His Memory Eternal!
Assuming RSI is something that is possible as you envision it in the near term. I think that it will happen at some point, but I think we could still be a long way off. I don't think anyone can truthfully say that it is right around the corner.
Especially with Mistral taking a fraction of the investment of the big guys. They can maintain the position pretty comfortably just by staying within a standard deviation of the leaders.
> Its quite interesting to see that at least the early days of AI so far have not been a winner-take-all runaway acceleration game where catchup is impossible.
Mistral is also an European company. As we live in a time where the US regime is engaged in pyrrhic geopolitical tactics, it's good to know that it can't threaten to cut access to models during s period where everyone is rushing to incorporate them more and more in our life.
On a purely technical level, maybe? But in terms of actual revenue, is there really any chance of anyone catching the big labs?
Obviously, this is only a valid question if you don't believe that open weights are about to eat their lunch and their revenue is about to collapse, or they're running a super unprofitable ponzi scheme propped up by investor money that's about to collapse like a house of cards. I don't find those positions credible at all though.
If you do, then this question isn't really for you, as I'm more interested in thoughts from those who think that OpenAI and Anthropic in particular are about to be the largest companies on earth in a couple years. Could anyone catch them at that point?
> But in terms of actual revenue, is there really any chance of anyone catching the big labs?
I don't know about revenue, but I suspect multiple other labs are already beating OpenAI/Anthropic on profitability. Staying on the frontier is expensive, and it's hard to recoup those R&D costs when you have a bunch of other labs nipping at your heels.
If you concede the previous point, then the only way for OpenAI/Anthropic to keep growing long term is to swallow the whole economy (i.e. mass job replacemnt), and that's a bet I wouldn't take.
They are very unprofitable…? I don’t think that’s really in dispute. We haven’t yet seen a profitable frontier lab and model pricing remains fairly subsidized
Seems silly not to have predicted that 10 years ago. I feel like it's long been obvious that smarter models being available will mean way easier cheap synthetic data and access to tools that will speed up competitors as well as consumers.
I tried it a bit and I like it! It is very fast via openrouter (significantly better than Kimi K3) on webui. Very verbose and starts to forget instructions after awhile it seems, but it gave me quite a lot of good info during a half an hour chat on C and embedded programming. I think I will keep urimg this as my main assistant for few weeks.
Man, a lot of this discussion sounds like people cheering for the last kid crossing the finish line.
Surely we want competition and Europe involved in that, but at this point I have grown used to either American labs smashing the frontier remarkably fast, or Chinese labs getting way, way closer than you would expect them to.
Mistral’s progress, regrettably, feels much slower. This model doesn’t knock anybody’s socks off. The model is (and I hate to be this harsh) mediocre, and this mediocrity has also arrived months late.
> Man, a lot of this discussion sounds like people cheering for the last kid crossing the finish line.
Sometimes it's ok to cheer for the last kid crossing the finish line because they're actually running a totally different race, and winning might look completely different.
When I look at what Mistral does vs other organizations I'm impressed:
They aren't profitable yet, but they're a lot closer than most and they're doing a hell of a lot with very little.
Pointless racing story:
I was in high school track with a really tough guy who was just not a runner. We went to a pretty messed up high school and if you screwed around in track practice sometimes the coach would make you run a crap race at the next meet, like steeplechase or hurdles. Well this guy and a few others screwed up and coach made them all run hurdles at a meet.
He hooked every single one and fell on his face. Every time he got back up and kept on running. By the time he hit the finish line his knees were bleeding halfway down to his ankles. We cheered like hell and he was smiling ear to ear.
I read that site quite differently from you. You seem to be analyzing absolute differences but ROI is really about ratio of spending to revenue.
It looks like Mistral is middle of the pack, behind Anthropic and ahead of OpenAI on that front. All of those labs are way "ahead" of the cloud providers, but those providers are building infrastructure, not just training models, so it's not apples to apples.
> This is a pretty grim prognosis for European AI.
I think it's an incomplete read. What's the point in competing for a sizeable percentage of your funding when the finish line is incrementally being moved each month? Better spend it on leapfrogs which they seem to have done.
Meanwhile Mistral have a natural ace in their pocket with respect to regulation in the form of CADA and the Cloud Sovereignty Framework. I can't think of another company that would qualify as SOV-3 under that regime
I don't think it will matter in a year or so. We are clearly topping out on useful intelligence for an increasing amount of tasks, as demonstrated by more and more models reaching the "useful" barrier.
This barrier is not going to start moving dramatically. It will simply be mostly satisfied for most work we do. Mistral is going to get there, soonish, long before the economy takes an entirely different shape (in so far that even happens).
There will be super human intelligence tasks, tasks truly constrained by intelligence for quite a while. Those will be few and far between, relatively speaking. Mistral will have plenty of opportunity to capture the other stuff, with a fraction of the resources required that it took the frontier labs to get there first.
People were extremely dismissive of chinese models until recently. They went from 1 year behind frontier to 6 months behind frontier to 3 months behind frontier extremely fast.
Important to point out that these 'X months behind frontier' really refer to the public frontier, and not the actual frontier, which private companies are free to protect indefinitely. Perhaps open models are in actuality 18 months behind the actual frontier - how would any of us know?
Curious, why do you say the model is mediocre? I haven't tried it, so I can't pass any judgement... I've learned to distrust benchmark rankings. Are benchmarks and Artificial Analysis the yardstick you're using?
> Curious, why do you say the model is mediocre? I haven't tried it, so I can't pass any judgement.
It's nowhere mediocre.
It's toes-to-toes with GLM-5.3 which is one of the best Open Weight model available (With Kimi K3) for general reasoning.
I just runned it on code reviews right now and it was able to catch some thread safety issue than DeepSeek-4.1 didn't. And DeepSeek-4.1 is by no means a bad model.
Doesn't matter, it's excellent, and it's European with European inference, which solves the pains of all my clients trying to build data lakes and processes on top of it.
Nobody in the real world cares about minor benchmark differences in money losing coding agents.
And nobody in the real world is giving Altman or Musk their data.
I wish you were true - but I still meet a lot of people handling sensitive data and using free or cheap version of ChatGPT or Claude with their customer data.
I think we will see some horror stories come out with data leak in the next years.
First few models will always be slow improving and worse. The way to improvement is working your way through a gajillion evals [1], finding bugs, gaps, and curating training data (this part involves human design as well as raw inference compute) to fix it. This is very time intensive and can't easily be "done once and then everyone has lesser work to do" since every model is different. Well, one way to accelerate it is to simply have more compute, which mostly openai and anthropic have[2].
This is mistrals first 1T-scale model and I expect the 4th or 5th generation to be close to the best for many purposes.
[1] These evals differ from the public ones like terminal-bench, are sometimes model-specific, need real, diverse usage to actually create, and are held secretly since quality of eval is the first driver behind the next step improvement of a model.
[2] It is not close. This model was trained on less than 4k GPUs, whereas astra used north of 100k GPUs.
Mistral is not that new a player though. How can we give them this much grace when other players like xAI have done more in even less time? I don't think coddling Mistral helps them.
And to the point of scale and training cluster, so what? Not only do Chinese labs have smaller clusters with less empowered GPUs, compute is Mistral's responsibility. You can't take away from other labs just because they fulfill that responsibility better.
Those are comments from Europe. The US is waking up now and I expect them to be much harsher.
I really want them to win as that's our last horse in the AI race, but ~200 research-oriented devs out of 1800 employees? I believe they agree it's pretty doomed and have pivoted.
It's a bit below DeepSeek 4.1 Flash at about twice the size. For a model that was supposed to come out a few months ago this is pretty good. Mistral catching up to the chinese open-weight models is great news. Excited to see how they will build on that!
Very nice progress. Also I respect putting Kimi on those charts. Regardless of if they are beating the Pareto frontier (not now), model diversity is a good thing for humanity — I’m hopeful for the team to keep increasing their gains.
I'm literally English and I feel the need to defend the French here, they're among the most successful military powers in the world historically speaking. It's hardly fair to judge a thousand years of French military history by the outcome of one war where they didn't do well.
I’d rather have one European AI lab like Mistral, with the financial firepower and compute to pretrain its own large models, than 20 weak labs that fine-tune Chinese models.
AI needs big bucks, and Mistral is Europe’s Anthropic.
Ah yes. Let’s thank the might US for providing pesky Europeans capital a fistful of dollars. But maybe let’s do it after Americans thank for Russian, Arabian, Chinese, European capital and workforce. After all this is what “Made in the USA” means.
The US has made sure that Mistral has a large market in the EU by temporarily preventing non citizens from accessing Fable.
A lot of European companies now want a model the US can't cut off, but also lack trust in Chinese models.
Some of these will self host Mistral but most will pay them by the token. It's not going to be a huge market or a huge margin within that market but probably it'll be enough.
This is the exact mentality that makes the EU fall behind. If you don't want to invest in something until it makes a profit, you don't get the benefits of being a pioneer.
> Europe needs profitable AI companies, not money pits.
AI is a strategic technology with obvious national security implications. EU should invest in its development whether it's currently profitable or not.
there are no profitable AI companies at the moment ... This is the exact problem of European startups, trying to make them profitable from day 1 while American counterparts (and Chinese) keep bleeding money for years. Europe will never have a Tesla, a Google or an Amazon with that mindset.
It's a joke, chonky is used to refer to fat cats, and there was a joke meme over the summer that Mistral are going to release a new model, Le Chaton Fat (chaton is kitten in French). The name is a nod to the memes.
The main thing I always get away from the comparison tables of these "big" models, is how well Deepseek v4.1 Flash performs. While still being the cheapest model by a long shot.
Beyond benchmarks, does it in day-to-day? Have always struggled to get competitive performance out of any Deepseek release going back to V3 vs Z.AI and Moonshot models. Maybe I really suck at whatever is needed to make DS models fly, but even tailoring my suite hasn’t gotten me far when I tried with V4 Pro. Happy for anyone who is able to leverage their models well, wish I’d be able to crack how to leverage them.
Will say their research is some of the best reads in the industry and I could not care less about their model release cadence as long as papers keep coming.
That said I love that they don't seem to restrict cyber capabilities to any degree and even lean into it.
If the best model for cyber attacks is open for everyone to use it just makes us all safer I think. Of course you then also HAVE to use it or otherwise you're vulnerable, which is a great distribution play.
I ran it on my trivia game Redactle which features a redacted Wiki article. Mistral Large 4 is not very good. It can sometimes solve a game with ~40 guesses whereas the top models like Gemini 3.8 Flash or Grok 4.7 can one shot most puzzles. My benchmark here aligns with AA Omniscience. I also have a version where the text is rewritten to detect over fitting to exact wiki text which changes the scores but not the leaderboard order.
Gemini models are summarizing wikipedia articles all day (when being used for Google's ai answer), can we draw some conclusion from this, did they train it extra well on wikipedia content?
>If the best model for cyber attacks is open for everyone to use it just makes us all safer
Issue is..
I don't believe for an instant that any of us, including US citizens, get access to the best models for cyber that the US has. I think any adversary would have to assume the models in use by the US side are unreleased.
US is not the only one dealing under the table by the way, I also think everyone should take China having unreleased models as an operating assumption at this point.
So Mistral is the best that the public gets access to. And that's if it's even the best? Benchmarks and pragmatic work have often been shown to be two radically different things in this industry.
Seems hard to get customers at that price range when you're competing with open source models that are 1/2 - 1/3 the price but with similar capabilities.
People usually buy the cheapest, like Deepseek or GLM or they spend on Anthropic/OpenAI subs.. Are these models in the middle getting any users?
On a side note, I wonder if this was the popular free Space Bunny model that left openrouter yesterday.
If you are a EU company worried about your data then mistral is your only option. Think about EU military companies. They can't use US and Chinese models.
I think you misunderstand how data processing works in a LLM. You absolutely can download the weights of a chinese model and run it on hardware you control.
I'm at a loss as to what to do now. I've been wanting to support Mistral for so long. I struggled on with Mistral Medium 3.5 for longer than I should have (although I also think it taught me some valuable process lessons).
Recently I switched to the Mistral hosted GLM-5.3, this worked very well and powered through a tonne of work. Unfortunately, I also completely maxed out two subscriptions within the space of six days this month. One can't stack subscriptions with Mistral, so I'd have to register a third account for another subscription, which will be annoying with changing API keys all the time. Sure I can switch to pay-as-you-go API, but that adds up really fast. The Mistral dashboard shows that a Vibe CLI monthly subscription for €18.44 actually provides €255 worth of API use (apparently, and I tried to check this with Support but it seems like they were intentionally vague).
After maxing out my Mistral subs this morning, I dropped $10 on Xiaomi to try MiMo-2.6. So far so good, seem to have done a lot of work for the $2.85 I've spent, and Xiaomi prices are still much better than the Mistral introductory offer for Le Chonk.
Not sure where to jump.
Edit: Not being able to stack subs is my biggest gripe with Mistral. I'd probably pay them $100 per month (5 subs worth), but I'm not going to switch to the pay-as-you-go API and burn much more money for the same amount of tokens. Instead, I've taken that extra money elsewhere. If they just allowed one to keep topping up subscriptions on the same account it'd be grand. Or even a bigger single subscription. Make a $100 tier with five times the capacity.
Yes, but I still don't know where the truth is.... The Mistral dashboard shows a drawn down on €255 worth of "credit" when you have a Pro subscription, which costs €18.44. There's no way I'm going to use the pay-as-you-go API if it means I'd be burning €500 in the next six days, as I just ostensibly have in the last 6.
I appreciate I can set a monthly spending limit for the pay-as-you-go API, but I'm just not willing to find out how far €100 will go, when I know it will go further elsewhere.
I just wish they had that €100 subscription tier, for the equivalent of €1k pay-as-you-go use.
Curious that it scores higher than opus 5.5 in cybersecurity because the closed models refuse to comply. I wonder if that means it's more susceptible to offensive uses.
You are quoting the discount pricing. It is 50% off for the next two weeks. The blog post has the real pricing up front in the card on the right side: https://mistral.ai/news/mistral-large-4/
After that, it will be much closer to GLM 5.3, but you can also get 5.3 in their API! I dont see people really talking about that.
I've been dreaming of this for a simple reason: the french prose combined with GLM 5.3 reasoning capabilities.
GLM 5.3 is incredible because for the first time with an open-source model, it feels.. enough. I don't need much anymore, this model is great in everything. Except a thing : speaking french.
If the benchmarks are true, I'd be glad to switch entirely to Mistral.
Excited to try this. The low costs v. benchmarks alone here are worth a serious test. K3 has been my daily driver for a month or two now and it's dramatically reduced token spend (while not having much of a negative impact on productivity).
This was the era of the AI race I was waiting for.
Haven't tried it yet. It looks like sol is a hair more expensive on input, a hair less expensive on output ($2/in, $10/out per 1m, K3 = $0.82/in, $13/out per 1m).
Maybe I'm missing something. Doesn't seem super impressive to me. A proprietary model with performance comparable to GPT-6 Luna and Deepseek 4.1 Flash, but at a higher price than either. The main selling point is that it's made in Europe... not very compelling, globally. I suppose maybe there is some niche where European-hosted open-weight models aren't enough to satisfy some EU regulation, where only the use of European-trained models is in compliance, but as a non-European I have no idea what that niche would be.
Side note: Wish this thread was more focused on talking about the model instead of debating about China and America. Whatever happened to staying on-topic?
You yourself also kind of pointed out why the discussion about USA and China is not off-topic. The niche Mistral wants to fill (afaik) is that it's European. I, as a European am really happy that they made such progress in so much worse (financial) conditions. I guess it's mostly geopolitics.
All things considered I'm more inclained to pay for EU-based AI in the end (supporting local company and most likely being more aligned with EU regulations…)
Nice! Once they make it available through their API I will be happy to support them. My local Qwen3.8 27B is serving me well, but I miss the speed and concurrency that comes with subscriptions, and I am not currently paying for any.
Sans context, I really like the Anthropic and Claude's faux-academic, minimal-ish, intellectual-ish branding.
(Especially the way it looked ~12 months ago -- it's gotten more cluttered since then. Perhaps unavoidably, as the breadth of their offerings has grown)
But over time it's begun to feel like unsettling cognitive dissonance as their ambitions grow and the stuff to worry about has piled up.
I know the whole “AI logos look like buttholes” thing is a joke for most people, but I really speaks to the pornification of our society. I would never have thought “butthole” looking at any of their logos.
I think there's such a thing as throwing too many marketing and sales people and too much polishing and "refinement" at something. I see in new product announcements from Microsoft as well. It's like seeing someone try too hard to impress you.
I do my best to avoid any AI marketing because I extremely despise it.
I just haven't quite decided on my new profession, yet, but it's either going to be with plants or with animals.
Well, with this and Beam people are going to have to stop saying that western open models are dead. This is great news. Anthropic and OpenAI may have a bit more knowledge, talent and compute but they don't have a monopoly.
Investors looking for them to make monopoly profits are going to be disappointed The premium they can extract from consumers for their models will be capped. Tokens are likely to remain close to the cost of compute, a cost which is high but falling fast.
Also, yay Europe! Although the comparison between this and mimo 2.6 is not flattering...
Still second most expensive open-weight model. I don't care about cybersecurity index.
And still can't beat Chinese models but good to see European in the game.
> The RL run behind this preview is still in flight and shows no sign of saturation -- we will release a final version before the end of the month along with the weights of the model.
Give them time. ML 4.0 was just pretrained. Mistral will certainly use it as the base for distillation and RL for smaller, better, more efficient iterations, just as the competition does.
I think it's actually a wink and a nod towards the social media meme of "Le Chaton Fat," a fictional model that is jokingly attributed to Mistral, usually with century-defining benchmarks and unfathomable size.
I tried Mistral's last 3 large models, devstral and mistrallarge3 and the numbers were not even benchmaxxed, but just false. the models were so weak and garbage. Let's hope they are telling the truth this time around, we need more alternatives. There mistral-small and original MistralLarge models were awesome, hoping they are back!
Excited to test this on my benchmark[1], but I'm guessing it won't fare better than MiMo v2.6 Pro which is currently the best open-weight model as far as I'm concerned.
Why isn't this on openrouter yet? Is there a better router that gets models much quicker?
I tried a quick abbreviated kimbench on their playground before bothering to do the whole thing.
Maybe I didn't really select mistral 4? Either way, failed completely on question 1 and the next 2 questions were completely off base too. I didn't bother to finish.
Is the reason for the massive gains in certain benchmarks due to distillation from the other lead models hence the slightly "under" pattern seen in the comparison charts?
I tried testing it, but reasoning effort parameter doesn't seem to be there and output is sort of broken because it reasons directly in the output tokens...
All things considered I'm more inclained to pay for EU-based AI in the end (supporting local company and most likely being more aligned with EU regulations…)
What experience have people had with Mistral models for cyber research? are they as constrainted as Anthropic models? I use claude day2day but have the need for a model with fewer guardrails to pentest our own APIs.
They explicitly lean in to cyber work, and they appear to be very permissive from their marketing:
> This is particularly important in cybersecurity, where provider-level refusals can block legitimate vulnerability research and incident response, and where losing access to a capability mid-incident can itself become a critical security risk. ML4 pairs top-tier cyber performance with open weights and self-deployment, giving organizations both the capability and the autonomy to run advanced security work under their own policies.
> That top score reflects a practical advantage. Several leading closed models, including Claude Opus 5.5 and GPT-6 Astra, score near zero on the same test because they refuse to perform the task. Yet defending software often starts with proving that a flaw is real, exactly the kind of work safety filters in closed models can block. This matters even more as threat actors increasingly jailbreak those same models to support offensive cyber activity
Seems like they've finally made a genuinely competitive model since the original LLM craze, congrats to them! Glad to see some diversification in open weights providers.
Does it have the ability to capture the market like OpenAi or Anthropic ? My point is, Regular/Average users of AI do not really care about benchmarking. Marketing decides which company makes it to the phones or PCs.
> Marketing decides which company makes it to the phones or PCs.
Right now a big part of LLM market is people using it for professional software development. Most of these users probably care about the quality of the model and also notice it during daily work.
For normal consumers, shure it doesn't matter. In the end the ai summary of google will probably be the most used as they are already exposed to it anyways.
dont take it personally, i just dont understand why to release a model that is not showing new strong capabilities, why would anybody use this model and not Claude Opus.
Without exaggeration, given a choice between models, I would pay for Mistral's model over Anthropic's based on the name alone, completely ignoring features or other technical considerations. The name is playful and is such a refreshing contrast to Anthropic's (and OpenAI's) doomsaying, scaremongering, and god-posturing.
This sure is nice, but I've had less than satisfactory results with GLM5.3. I'd like Mistral to compete with Qwen3.8-Flash-Next a 120B class model that IMO is the first model that I can use for serious coding while running it locally.
I estimate it's coding ability on par with opus 4.6 (but opus definitely beats it on factual knowledge) Still it's a genuinely useful model, when everything else except Anthropic's models (and for only 3 weeks after it came out Google's Gemini 3 pro, before it got merged) are not to me.
I really hate this open weight but closed science approach. These companies just take from academics and all the Chinese companies that are doing good science, but without understanding the recipe it makes it hard to know where the failure points will be until your agent accidentally commits a crime.
if it's not available yet why have a 'try it today' header at all?
> "Try it today"
>
> There is still more to come. As we work toward releasing the weights, we will share further details on the model architecture, additional benchmarks, and our post-training methodology.
If you scroll down, most of the charts on that page are sorted s.t. Mistral's bar is right next to the worst competitor model, while the best competitor model's bar is positioned on the opposite side.
If one were to be cynical one could say that it's intentionally making Mistral's result look better than it actually is by making it harder to compare the bar heights.
Yes broad consensus had developed in the ten minutes between announcement and you asking if consensus had developed, and I'm excited to report that it consensed in the affirmative -- it IS Space Bunny Alpha!
Bar charts should start at zero. If they don't start at zero, there should be a clear visual indicator that the chart has been trimmed without having to read the axis labels. I hate that this has to be repeated so often that it has become a cliché.
A bit disappointing to see it still lagging behind Chinese open models. Those Chinese models are pushing proprietary models to raise the bar, but we need equally strong non-Chinese open models to challenge the Chinese ones in turn.
I mean no disrespect but these are terrible numbers or am I missing something? It seems like Mistral continues to only be relevant for people that want a model trained in Europe. Too bad.
1T parameters -- ugh, open models keep getting bigger and bigger! Running them at home is getting ever more unattainable, especially for those of us with bandwidth-poor hardware like Apple silicon -- please continue releasing smaller models, too!
Too bad this got marked as a dupe, as it actually has benchmark info unlike the other page which is just docs.
The weird thing is how worse they are at things like coding than other open weights. You'd expect them to at least distill coding from other open weights to match them.
Pretty impressive. I genuinely wonder how Mistral hires talent when their salaries are so terrible. Guess there aren't many better places to work in Europe.
Non-American maybe, but non-Chinese is likely impractical. We might not want their APIs, but we can't compete on energy and labour for training, even if it means buying licenses to host the weights (which CADA encourages). From China's perspective, what's the case for baking weights they can't sell? Some Chinese models already don't even have the Taiwan politics stuff baked in, those filters are only in their APIs.
I realised after writing this, the EU as is uncomfortably often the case, may be the real forcing function for what happens with US policy irrespective of the media campaigns we're presently seeing. Here's hoping for a steady trickle of stale ChatGPT weights leaking from EU infra providers in the long term.
Mistral is the only major non-Chinese contender in the AI race releasing open-weight models. I’d rather trust that French weights haven’t been backdoored than Chinese ones. Heck, I’d even trust Mistral more than Anthropic for sensitive work.
You have a number of other European providers of open-weight LLMs, such as Bielik and Aleph Alpha. They are not trying to compete at the frontier, but they sometimes develop original architectures. I am also a big fan of PleIAs' research in this regard, have a look at their Baguettotron
Mistral wont win the AI race because of the model names. I wont bother an arrogant Parisian hipster with my insecure prompts who then plays with his moustache and responds with a judgmental "pfff"
Does Mistral ever advance the state of the art on any dimension?
And if not, why do they exist?
Update: The number of people advocating not innovating is wild. There is no reason why Mistral cannot innovate in ML, they explicitly choose not to. My point is that, given that choice, they should spend their GPU hours differently.
"Sovereign AI" is a joke, there is no substantive difference between a post-trained open weight model from an American or Chinese company and what Mistral is doing today, beyond spending 80% of their GPU hours reproducing a last-gen model's pretraining.
Why would the Europeans want a high-quality open source model that isn't either owned by American Trillion-dollar companies or created by the Chinese? You really can't think of any reasons?
I don't know what definition of open weights you are using, but the "open" part is what allows you to use it without being beholden to American Trillion-dollar companies or the Chinese.
Mistral could do exactly what Cursor does, and post-train the model to meet the needs of Europeans.
That's a far more efficient (and useful!) use of GPUs than yet another pretrain on an outdated model backbone.
Just like Microsoft and most companies around PCs, their existence isn't about pushing any boundaries, but seemingly targeting deploying what's already been "invented" into the corporate world and bigger companies. Not "wrong", just different.
I try to stay up on AI models, but I've given up on Mistral. I have tried using it too many times and it's the worst out of major models. Even Kimi and DeepSeek are miles ahead.
Seems like it's another European company that is only alive through government financing.
I understand wanting a home grown industry, but with AI models, just abliterate a DeepSeek model.
Total misunderstanding of the industry. Europe needs hardware, not a model that needs to be replaced monthly.
> Total misunderstanding of the industry. Europe needs hardware
With all due respect, I have no idea what you are trying to say with this comment and have no clue how "hardware" would turn Europe's tides. You explained nothing.
Do they need training hardware? Inference hardware? ASICs or GPGPUs? Edge hardware? Agent hosts? Faster cores, or wider ones? Taller systems, or more parallel ones?
Your vagueness completely undermines the authority that your criticism relies on.
Surprisingly it only supports reasoning "none" or reasoning "high".
That setting didn't seem to make any real difference - it added a tiny bit of thinking trace and high actually produced less output tokens than none.
The high bicycle frame is better then the none one though.
Pelicans: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
(Definitely the best I've seen from any Mistral model: https://simonwillison.net/tags/pelican-riding-a-bicycle+mist... )
This is such a pristine pelican. Let me say it here first folks, AGI is here.
> AGI is here
If AGI is "Attractions to Get Investments" then yes, it's happening
3 replies →
Avian Graphics Intelligence achieved!
Pelican benchmark is saturated, anyway. :-)
Nah, its beak is still too small to hold a capybara.
1 reply →
This must be a joke, clearly pelican drawing in svg would be in its training data after multiple years of it hitting HN front page.
I wonder when will we see a photorealistic pelican on a bicycle in SVG format.
> AGI is here
Far from it. This shows a strong ability to generate an image known to be frequently used as a model test. This isn't a measure of thought.
10 replies →
The benchmarks against Opus 5.5 and GPT-6.1 Sol look pretty good for 3D generation: https://x.com/atomic_chat_hq/status/2107516529608700383
I wonder if the other models were worse than usual for that particular video because they "dislike" making an advertisement for another company's model. A test with a generic video might be more meaningful.
Got an example hosted on a platform that isn't owned by Twitter's loser owner?
I think it's curious that it has so many shared elements with the latest Astra pelicans: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
- Sun on top right
- Cloud on top left
- Three "speed lines"
- Two feathers on top of the head
- Eye rendered as a black circle with smaller white circle inside
I wonder if the pelican benchmark is converging across models due to past results being used in training.
These are pretty much what a human would draw. Sun rises from the east. A cloud makes the background “sky”. Three lines is the minimum to interpret as movement. Two feathers is standard on every cartoon and illustration.
This comes up every thread. I think we've all noticed how similar they are becoming.
I would guess that they've definitely been trained on previous results, as they obviously share way too many traits at this point to be totally random. That said, I don't think we're seeing any signs of pelicanmaxxing yet from the providers, so it's still a useful (or at least fun) benchmark.
Once all the models produce pristine, elaborate pelicans riding perfectly drawn bicycles, then it'll be time to move on to pigs driving a racecar or something.
Why are pelicans almost identical across different models?
I recently was testing something, I asked some models to provide me a single random word:
I have enough projects, I think some benchmark/dashboard showing kinship based on these kind of queries could be very interesting to watch and insightful when new models come out.
22 replies →
The benchmark is saturated. Frontier models are tested with an armadillo in fishnet tights jaywalking on Mars.
9 replies →
I think they're still visually pretty different. The most common shared details are:
- Pelican cycling to the right - that's been discussed at length, images of bicycles online always show that side of the bike because that's where the chain is.
- Bicycle is usually red. No idea! Red ones go faster?
Also, why are they almost always riding from let to right?
2 replies →
They aren't. You aren't looking closely. For example, the first image does not have the frame of the bike in the correct shape even.
Everyone is stealing from everyone else.
Because it's a terrible benchmark
That's one heck of a pelican!
Not bad! I like how it got the motion lines on the correct side. IIRC, many of the other ones you've posted have the motion lines on both sides of the pelican
The difference between high and none is the bicycle.
The bicycle looks significantly better in high. And feet and hands are actually where they should be. The road looks worse, though. No flowers either. And in neither is the pelican sitting on the saddle, but I can understand it's hard for a pelican to ride a bicycle properly.
Now what would have been cool is if Mistral on high reasoning had realised that pelicans are the wrong proportion to ride a bicycle, and had designed a bicycle more suited to pelicans. Let me know if any model ever manages that.
1 reply →
Two pelicans, one shape. The difference is load-bearing, and that's the big unlock.
That beak is CHONKY.
I tried testing it, but reasoning effort indeed seems to be broken somehow.
High one is actually much better. The feet connect to the pedals, the wheels don't have a hub cap, although it looks like the pelican is wearing the seat, it's in a relatively proper position etc.
Both are riding on the left side of the path for some reason.
This is entirely stochasticity. The entire reasoning trace was:
> Create a cartoon pelican riding a bicycle. Need SVG only output.
[dead]
I have a question, and perhaps some of the AI/ML infrastructure experts here could answer: Mistral says, "ML4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own datacenters in Europe."
If a 1T params model trained on ~4k NVIDIA GB GPUs could almost match the performance of Kimi's K3 (which is on par with top closed source models of OpenAI/Anthropic) while beating/exceeding other leading SOTA models from top Chinese labs, what are we (in the US) even building these super massive data centers for? Just to churn through more backpropagation reps more quickly?
SpaceXAI's Colossus supercluster in Memphis and Colossus 2 in Memphis/Mississippi (Southaven) are supposed to run into hundreds of thousands to a million GPUs. MSFT's Fairwater GPUs are supposed to have hundreds of thousands as well. So, 3800 GB GPUs are an absolute drop in the bucket. I don't understand the strategy of hyperscalers here, especially with edge inference hardware only getting better from here on (Apple, and all).
Distillation explains some of the advances, but doesn't that mean hyperscalers have a ton of deadweight wrt GPUs sitting on their balance sheets? Will all these GPUs be used for inference once a SOTA model's training checkpoint/batch is done? It's bonkers to me.
> Will all these GPUs be used for inference once a SOTA model's training checkpoint/batch is done?
I have no real data to back this up, but that has always been my assumption.
Claude says that K3 can be assumed to have required 10-100M GPU hours. If you have 100k GPUs that would mean like 6 weeks of training. 100k GPU's can serve 3-30 trillion tokens of K3 per day. Google apparently serves ≈100 trillion per day [0].
The big labs probably want to have capacity to fairly quickly train / post train different SOTA models continuously + being able to serve peak inference demand in valuable markets (US daytime?).
[0]: https://blog.google/innovation-and-ai/sundar-pichai-io-2026
It's unclear how big a role distillation plays, but it may be a big one. There's also a law of diminishing returns. To get a meaningful increase in model quality you seemingly need an exponentially larger model. And most people don't think K3 is actually on par with top closed models.
Google says 1 million blackwell gpus are being delivered monthly I'm really curious if it's companies just hoarding chips/memory/servers awaiting to be deployed in data centers not ready yet for months or years, or everything built is actually deployed upon delivery. Plus google and amazon have there own chips in the mix.
For inference (Anthropic and OAI are B2C on top of B2B)
To train much larger models. It is quite possible that 10T-100T models be on the horizon
Today's data centers are being built for yesterday's inference need. There's a persistent cult belief that ai hasnt found a niche or that companies havent proven utility or use cases or whatever. the demand for ai (internal to hyperscaler, and external for everyone else) simply dwarfs what is available.
Both Anthropic and OpenAI have been having major load issues though. Up until yesterday, OpenAI was serving at only 30t/s per default.
1 reply →
Impressive vision benchmarking. If the vision model is truly as good as astra, that would make it best in the world.
Also strong on cyber benchmarks (better than all chinese models), so this is a good defender model.
Lots of people shitting of Mistral for no reason imo. These are pretty good numbers across the board. Definitely good enough to use as a daily driver over other llms, if you have moral qualms with the others. For certain use cases, like cyber security, this may be the go to model.
I like to make fun of europe, but there's lots for mistral to be proud about in this release imo.
> Impressive vision benchmarking. If the vision model is truly as good as astra, that would make it best in the world.
Sure, it leads in one unpopular benchmark with internal numbers.
But look at the third benchmark. It only has Sol 5.6 from OpenAI and it shows a 12.8. I don't know this benchmark but AA's GDP.pdf listing has Sol 5.6 Max at a 27, even non-reasoning beats the 12.8. It's extremely weird to cherry pick Sol 5.6 and then also lie about the published third party benchmark score. I'd love AI companies to stop lying about this stuff.
AA - https://artificialanalysis.ai/evaluations/gdp-pdf?models=cla... Mistral announcement - https://mistral.ai/_astro/multimodal-benchmarks---gdp.pdf---...
I'm still embedded in the OpenAI ecosystem, but man, when I try out Mistral, it is snappy! Like basically instant responses that seem faster and better in quality than ChatGPT on instant mode. I'm impressed.
Curious, what do you like to make fun of Europe about?
The paradox between supporting consumer rights with actions like universal usb-c adoption, but also complete elimination of any privacy rights at all. Like the surveillance state is insane. No e2ee chats, backdoors in everything.
66 replies →
Like all the fancy jackets, tight pants and swords. That's hilarious
-- My imagination of Europe when I was 15 living in Ohio
Not the OP but it often appears that there's a hostility to tech.
3 replies →
Where to start?
Proud EU citizen here, but boy/gale/other, I have plenty to make fun of ;-)
https://m.youtube.com/watch?v=gGlpBuW6ZFc
This is a good summary.
2 replies →
The empty posturing.
The faux unity.
That they think their politicians are better than the United States but have populist demagogues, right wingers winning, corruption scandals, freedom of speech restrictions, and unbelievably bad spending problems.
Still waiting for France to send troops to Ukraine or ask the United States to get rid of their European military bases.
2 replies →
Americans just make fun of Europe for no reason, it is what it is
3 replies →
Europe has spent the last twenty years in stagnation - generating about half as much wealth and technology as you'd expect for its size and advanced economy. Simultaneously, Europeans are notoriously arrogant. It's a mockable combination.
Edit: unfortunately, it's a question that voting does not really permit you to answer on this website.
71 replies →
> Curious, what do you like to make fun of Europe about?
Well I'm european and... That Switzerland (Europe but not EU) has more companies in the Top 70 by market cap than the entire EU (Switzerland has two, the EU only has ASML) is kinda something that warrants making fun of.
That the biggest European software company is SAP, in 71th position is both sad and tragic: it shows how lame and irrelevant Europe is when it comes to software.
So Europe is nowhere in software and friggin nowhere in hardware: sure it's got ASML but ASML now has officially... Zero customer in Europe. Zero is not much.
Then Japan is at least trying to come back into the game with nano imprint litography. Europe is betting it all on AMSL (which anyway is majoritarily US-owned).
So software: nothing. Hardware: nothing besides ASML.
Overall the EU has six companies in the Top 100 by market cap and they're all, besides ASML, near the bottom of the Top 100.
We could also maybe make a bit fun of how the EU destroyed it's car industry (the main industry in Germany, which is the biggest economy of the union) by handing it all to chinese EVs?
Or what about the US warning the EU, years ago, to not become entirely dependent on Russia for energy? And EU not listening and then seeing its energy price skyrocket when the proverbial shit hit the fan? (Russia attacking Ukraine)
And we could, also, at least make a bit of fun of entire streets in cities like Paris and Brussels that used to have luxury shops and fancy restaurants that are all turned into places selling cheap kebabs? What a great success: I'm sure this one makes the komrades happy. It projects an image of grandeur and success: kebabs.
Or the constant attacks on free speech in the EU. Or the surveillance apparatus that's being put into place.
And let's not forget: there were promises made to Russia to never grow the EU to the east. Then the EU started exciting Russia by saying they'd incorporate Ukraine into the EU: I'm not against that but doing that did trigger a war. And now suddenly the EU is waking up and feeling all warmongering, wanting to dedicate a big percentage of its spending to weapons and tanks and missiles.
The warmongering tiny pet that the EU is is kinda laughable too.
At this point it's more like I don't know what is there left to not make fun of about my EU.
For what's going on is just sad, plain sad.
13 replies →
You mean besides the fact that all they do is complain without actually leading at anything?
In many ways the rest of the world would be better off if Europe was disconnected from the internet.
If the UK is in europe, you should make fun of the following: they don't have a first amendment. So to me, it's an authoritarian state preaching freedom.
22 replies →
Even if it's not the best model, it can be really important step in UE sovereignty. Trained in EU, inference in EU. I guess it will matter for some companies. Hope Mistral won't disappear for the next half year.
For sure, as an EU based company we will only use US suppliers for coding but never for fuctions in our own products. Mistral knows this.
But coding/development is way more important. That's where the IP goes straight into the next training dataset.
3 replies →
the greatest threat to Mistral is that there isn't a deep and large enough capital market in the EU to absorb the valuation step ups needed by the handful of EU sovereign growth investors to justify their existence.
you could say, that's perpetually a tomorrow problem, so long as they never go public, but that should illuminate for you: if everyone "knows this," it's inevitable that this so-called EU company, that didn't invent any of the AI, the hardware nor the training data, where their product is essentially more like a VPN provider than a frontier technology company, just lists and capitalizes in the US anyway.
Seems some AI they did invent: https://news.ycombinator.com/item?id=49243397
By your measure anyone unable to build EUV lithography machines without ASML help is doomed. EU could easily corner the market and shut them all off. No more NVIDIA.
Let’s grow the pie, shall we?!
1 reply →
Sovereignty over what though?
The last 3% of the stack? They’re essentially borrowing Chinese open source distillation/innovation off the US frontier running on US/Taiwanese designed chips and calling it EU made. This is better than nothing of course.
But is sovereignty really an end in itself? So the EU becomes IT independent…cool, then what? I mean the Yugo was sovereign, it didn’t do much good when the society itself failed to produce prosperity and collapsed.
It seems to me the EU is expending enormous effort on the appearance of “sovereignty” over what is ultimately…the last-mile consumer toilet paper purchasing conversation data…of an aging, increasingly irrelevant population on the political/economic stage.
Meanwhile domestically the entire economic model is failing and the “union” is getting shakey as its 2 biggest members turn more nationalist.
Maybe this “oh you cant compete but here’s a trophy for sovereignty” attitude isnt helping. Less clapping along with the EU bureaucracy’s latest make work projects like transitioning to a new Word processor. If we want the EU to succeed, more tough love is needed imo.
it's industrial policy, which is of course the minimum for being able to even think about having independent thoughs when it comes to international relations.
hard to stand up for EU (or even national) values if some dude in Washington has the keys for most of your military
having at least some in-house expertise is the first step.
1 reply →
Just ran this through our data analytics benchmark (I work at Plotly).
It's 10x cheaper than Mistral Medium 3.5 from April and goes from 58% to 74% correct. Definitely a generational shift.
It's not on the Pareto curve yet, but it's good enough for data analytics, and at this rate I suspect it'll be excellent in another few months.
Full write up: https://plotly.com/blog/mistral-large-4-plotly-data-analytic...
If I'm reading that plot correctly, qwen3.8-27b beats in accuracy and price?
A strong competitor in cybersecurity as an alternative to GLM-5.3 (Mistral reports 82% on CyberGym-E2E). Visual grounding is also impressive (42% on Dense 200 vs. 41% for GPT-6 Astra).
Otherwise, behind on the broader Pareto frontier, but not by much (Vals Index: 48.05% vs. GLM-5.3’s 53.51%; $13.78 vs. $7.25 per test). Many companies will prefer it over Chinease models.
That was about my conclusion as well: Slightly less than GLM 5.3 performance but made in Europe. So, maybe it answers Tiananmen Square questions correctly, and in French. All in all, a reasonable model, but not frontier.
> So, maybe it answers Tiananmen Square questions correctly
What would you consider a "correct" answer? I just asked deepseek-v4.1-flash asking it what happened (without mentioning the word "protest"; here are some excerpts of what it said:
> In April 1989, students in Beijing began demonstrations after the death of Hu Yaobang, a former Communist Party general secretary. The protests grew. [..] Estimates from other sources range from hundreds to several thousand deaths. [..] The Chinese government describes the events as a counter-revolutionary riot and says the military action was necessary to restore stability. It restricts public discussion of the events inside China. Many other governments, human rights organizations, and observers describe the events as a violent suppression of peaceful protests.
So, let's see... it calls it a "protest", mentions the number of deaths, and even mentions the censorship of the topic by the CCP.
6 replies →
"..answers Tiananmen square questions correctly.."- but lies about Ukraine, EU, and about you, americans..
I live in EU, use for my personal needs chinese models only, and don't plan to move to any of the ones allied with the Pentagon or its european counterparts.
5 replies →
Its quite interesting to see that at least the early days of AI so far have not been a winner-take-all runaway acceleration game where catchup is impossible.
I certainly wouldnt have predicted that 10 years ago.
Very glad to see Mistral still in the game even after some big stumbles with Large 3. I deeply hope that this model is 'good enough' that it becomes the European go-to, giving them the resources to keep the pace up.
I'm excited to try this out today.
I think a big part of that is the Chinese publishing the solution for everywhere hurdle in the road they've encountered in the form of a paper.
Deepseek essentially releases instruction manuals in paper form.
I think it might have accelerated things but on a much more basic level, there seems to be no real moat in synthesizing the world’s knowledge into LLMs.
45 replies →
My spend on DeepSeek is not much and I regularly top up my balance every month as my support for all the good work DeepSeek is doing for the open science.
158 replies →
Architectural/algorithmic tweaks do advance the efficiency frontier nicely. But raw intelligence mostly comes from data (not just its sheer quantity, but also how it's curated & cleansed) and the scaling law. The know-how about data curation doesn't seem to get published much, even among the open-weight labs, though.
1 reply →
I'm not sure I like this framing - so much of AI research has been academic, in the open, building on others people's work. Much less comp sci generally, math & philosophy, etc. The idea that rich companies can just build stuff in secret because they have resources is a fantasy.
So boring to see conversations moved over to Chinese models when that’s not even what we’re talking about here. This is about Mistral.
3 replies →
Also, the field moves fast, but slower than people do. Researchers and engineers switch companies every year or two, and the know-how walks out the door with them.
How it should be. Knowledge should not be copyrighted. The world will be a better place with such information democratized
There would be a lot of competition even without DeepSeek. Workers can freely exfiltrate trade secrets without noncompetes in California.
1 reply →
>instruction manuals in paper form
So the most common way to publish manuals?
3 replies →
> have not been a winner-take-all runaway acceleration game where catchup is impossible
From the Mistral site:
> ML4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own datacenters in Europe.
It is pretty capital intensive!
That cluster is literally orders of magnitude smaller than the compute pools used by Anthropic or OpenAI.
3 replies →
I’m pretty impressed that they managed to get that close to the frontier with such a small cluster!
1 reply →
According to Grok thats 7-10 MW. Tiny numbers.
To put that into context, the last wave of capacity SpaceXAI added 400-450 MW.
1 reply →
These cards are like $3k each? That's, what, $12M and you keep the hardware? Honestly doesn't seem too bad.
2 replies →
That’s kinda very small and light for modern trillion-param LLMs.
For sure people who don't grasp the difference between models, might be stuck in 'good enough' models.
But Opus 5.5/GPT is such a game changer in comparison to sooo many others, its still a moat for now.
You've got to think the comments about "good enough" are people who have not yet tried Opus 5.5. I haven't been this struck by a step change since 4.5/4.6. It's a bigger jump even than when Fable first arrived.
As for Mistral - I got really excited when they said Large 4 was focusing on being #1 in cybersecurity, because that's somewhere that they genuinely could edge out Anthropic & OpenAI. Have it actually solve problems, instead of Anthropic flagging "you tried to find a null pointer exception bug in your own code, we're now reporting you to the US government". But on the Mistral benchmarks I'm seeing, this looks very disappointing, but at least they haven't entirely given up. I genuinely thought Mistral had given up on new general models. They need to learn the bitter lesson all over again.
14 replies →
I agree that Opus and GPT are almsot surely better, but so many real users are nervous enough about giving Anthropic and OpenAI access to all of their internal information that they may be willing to stomach worse models if it gives them more security.
The real question is if this model is good enough that it can still accelerate work, and not be a hindrance to real work like older Mistral models often were.
If they can do that, they'll have customers.
8 replies →
No it is not. Only maybe for the noobs or vibe coders.
People who aren't afraid of rolling their sleeves into any code base? The difference is practically zero.
5 replies →
Please, give it another 6 months and they catch up. The American labs are currently trying everything they can to block others instead of advancing their models, trying to build an artificial moat. The American models are not that great, they are good, and they have a lot of agentic workflows in the back, but its basically a hardware limitation at this point. Once the HW makers catch up, and we can move away from the Nvidia monopoly, things will speed up quite a lot IMO.
10 replies →
It's an improvement, but game changer might be a bit of a stretch. If I lost access to Anthropic or OpenAI models tomorrow, I would be annoyed, but would reach for a slightly inferior model. Last year I wouldnt be able to say the same, and rhe challenge is that the moat is drying up fast. Whether its general improvements in model training by other competitors, or straight up distillation of SOTA models, the moat is shrinking and the available capital and spend for American model providers is going to dry up quickly as competing good enough models are adopted by more consumers.
It's especially the case as more non-Americans look to self hosted models and domestic cloud inference providers using open models that the US providers who are still leading the charge need to drastically drop their prices and find a path to profitability in order to maintain their lead and retain the advantage they had as AI turns into a commodity (which is happening faster than I think even the frontier labs initially predicted).
People say this exact thing every single time a new frontier model comes out.
I use Opus 5.5 at work.
I use MiMov2.6Pro, DeepSeekv4.1Flash, GLM5.3, Hy4, Qwen3.8 and KimiK3 at home. Opus5.5 is not a game changer.
2 replies →
> X is such a game changer
I hear this literally every other week about whatever the newest FoTM model is.
Unless you can provide concrete examples of things you can do with them that you simply couldn't do with last week's model, it's absolutely meaningless.
Let me know when the game changes are more than a month apart.
Less and less work requires a frontier model though.
My todo app generator does not need opus 5.5
Yeah, I agree with this. I think the "the models are good enough" narrative is a myth. I've heard it so many times over the last year, but the model number keeps changing...
There is no ceiling on what you can accomplish with more intelligence, so there will always be a market for the best models, and that market is likely to just keep growing. If Opus 13.5 can one-shot a profitable company or discover a new disease treatment or whatever you can think of that a swarm of relentless super-geniuses could accomplish, companies (and governments) will throw money at it.
I also think there will always be a market for many sub-frontier models that will continue to grow rapidly as well, because "good enough" is definitely a thing for a given task.
2 replies →
Despite Opus 5.5 got really bad the last days for me. Looks like they nerfed it again. This is extremely unreliable.
2 replies →
It's shaping up to be much more like a game of 'chicken' where each company tries to raise more cash without going bust... Ultimately the game of musical chairs is going to have to stop. In the US it looks like they are trying to get a government sanctioned truce in the form of regulation. That's what 'Pacing the frontier' means...
Even "runaway acceleration" isn't instantaneous. People imagine the singularity as something that happens almost instantaneously. But obviously it happens over time, and that time might be decades. It might still end up looking like a vertical line on a long-term graph.
If the singularity is defined as an AI sufficiently intelligent to improve itself independently, that AI is still limited by the resources required to do this improvement.
Am I reading this correctly?
This appears to be roughly as good as Sol 6.1 (which is quite good), considerably faster in terms of wall clock for complete tasks, and considerably cheaper (where Sol 6.1 is already good value - just really slow).
That seems too good to be true...
But I really hope it is true...
I can confirm that you're reading this incorrectly. There's a reason behind them only comparing it to open-source models released months ago. Here's a good aggregator: https://artificialanalysis.ai/#intelligence
1 reply →
It will become winner take all when AI companies manage to really get value from user logs.
Right now they don't even get good feedback from local sessions - I can see it make the same mistake two days running, and then months later when a new model comes out, presumably trained on my data, it still makes the same mistake.
law of diminishing returns? i.e. any reasonable frontier lab will have enough user logs...
1 reply →
Hear! Hear! I really want European models / AI labs to succeed.
I trust them and their populations to provide a more societal-friendly version of AI, putting pressure on the US tech oligarchy, while also providing democracy-friendly open models that I don't trust to happen with the Chinese labs.
What do you mean "good enough"? Did you mean "large enough"? ;)
Disclaimer: I'm not sure how much of an IYKYK factor applies to this joke.
We haven't reached RSI yet. Once any entity reaches RSI, the runway scenario will happen.
Truly, this is what the Lord's prophets have revealed to us! (Eliezer 11:52) Keep strong in your P(singularity), for when the Kingdom arrives, He shall judge us in His righteous glory, whether to eternal annihilation, or rebirth and life in His Memory Eternal!
Assuming RSI is something that is possible as you envision it in the near term. I think that it will happen at some point, but I think we could still be a long way off. I don't think anyone can truthfully say that it is right around the corner.
Hard to do that with a commodity that's easily replicated.
I strongly disagree with this "early days" framing.
AI is an idea 60 years old. We are on the 3rd or 4th generation of AI development. Three years into the current iteration of products.
This is not early days by any measure. LLMs are a result of a very, very mature research field.
Especially with Mistral taking a fraction of the investment of the big guys. They can maintain the position pretty comfortably just by staying within a standard deviation of the leaders.
Yes, so far the competitive dynamics feel more like cloud computing than web search.
> Its quite interesting to see that at least the early days of AI so far have not been a winner-take-all runaway acceleration game where catchup is impossible.
Mistral is also an European company. As we live in a time where the US regime is engaged in pyrrhic geopolitical tactics, it's good to know that it can't threaten to cut access to models during s period where everyone is rushing to incorporate them more and more in our life.
I think it's a mistaken belief that AI as we found it is the exponential runaway train.
So it makes sense, since all you need is compute, that there's a ceiling and specialization is going to be more valuable then some super AGI.
Especially since the worst people seem to be the ones who think they'll all run away with the bag.
I mean, Mistral is about 9-12 months behind here when you look at its overall benchmarks versus the models released around a year ago.
Sounds ok to me. Claude was fine at the start of the year, and now with Mistral you also get EU sovereignty? I'll take that.
It's a retrain of asian model.
How do you figure? I haven't met a single person who doesn't use Claude or Codex for programming in any serious way.
Then you dont know people working on highly sensitive info with stringent privancy concerns.
1 reply →
I'm happy I've never met you.
I don't think it is, and I think that is what will pop the bubble. All these companies have winner take all valuations, and that won't happen.
... unless they can legislate it, which is why they are flattering heads of state and scare mongering about dangerous AI.
On a purely technical level, maybe? But in terms of actual revenue, is there really any chance of anyone catching the big labs?
Obviously, this is only a valid question if you don't believe that open weights are about to eat their lunch and their revenue is about to collapse, or they're running a super unprofitable ponzi scheme propped up by investor money that's about to collapse like a house of cards. I don't find those positions credible at all though.
If you do, then this question isn't really for you, as I'm more interested in thoughts from those who think that OpenAI and Anthropic in particular are about to be the largest companies on earth in a couple years. Could anyone catch them at that point?
> But in terms of actual revenue, is there really any chance of anyone catching the big labs?
I don't know about revenue, but I suspect multiple other labs are already beating OpenAI/Anthropic on profitability. Staying on the frontier is expensive, and it's hard to recoup those R&D costs when you have a bunch of other labs nipping at your heels.
If you concede the previous point, then the only way for OpenAI/Anthropic to keep growing long term is to swallow the whole economy (i.e. mass job replacemnt), and that's a bet I wouldn't take.
2 replies →
The big lab revenue may not be catchable, but im not sure it needs to be.
If they can carve out a niche of industrial and governmental partners who rely on them for sovereignty reasons, it may be enough.
2 replies →
They are very unprofitable…? I don’t think that’s really in dispute. We haven’t yet seen a profitable frontier lab and model pricing remains fairly subsidized
No, because compute, not model ability, is the moat.
The second moat is convenience, which all the big labs make it (comparatively) easy to glide into their models.
1 reply →
[dead]
Seems silly not to have predicted that 10 years ago. I feel like it's long been obvious that smarter models being available will mean way easier cheap synthetic data and access to tools that will speed up competitors as well as consumers.
Has anyone tried Mistral for the coding task? How do you find it compare to Claude, Codex or the Chinese models?
I tried it a bit and I like it! It is very fast via openrouter (significantly better than Kimi K3) on webui. Very verbose and starts to forget instructions after awhile it seems, but it gave me quite a lot of good info during a half an hour chat on C and embedded programming. I think I will keep urimg this as my main assistant for few weeks.
Man, a lot of this discussion sounds like people cheering for the last kid crossing the finish line.
Surely we want competition and Europe involved in that, but at this point I have grown used to either American labs smashing the frontier remarkably fast, or Chinese labs getting way, way closer than you would expect them to.
Mistral’s progress, regrettably, feels much slower. This model doesn’t knock anybody’s socks off. The model is (and I hate to be this harsh) mediocre, and this mediocrity has also arrived months late.
This is a pretty grim prognosis for European AI.
> Man, a lot of this discussion sounds like people cheering for the last kid crossing the finish line.
Sometimes it's ok to cheer for the last kid crossing the finish line because they're actually running a totally different race, and winning might look completely different.
When I look at what Mistral does vs other organizations I'm impressed:
https://isaiprofitable.com/
They aren't profitable yet, but they're a lot closer than most and they're doing a hell of a lot with very little.
Pointless racing story:
I was in high school track with a really tough guy who was just not a runner. We went to a pretty messed up high school and if you screwed around in track practice sometimes the coach would make you run a crap race at the next meet, like steeplechase or hurdles. Well this guy and a few others screwed up and coach made them all run hurdles at a meet.
He hooked every single one and fell on his face. Every time he got back up and kept on running. By the time he hit the finish line his knees were bleeding halfway down to his ankles. We cheered like hell and he was smiling ear to ear.
Coach quit punishing us with races after that.
I read that site quite differently from you. You seem to be analyzing absolute differences but ROI is really about ratio of spending to revenue.
It looks like Mistral is middle of the pack, behind Anthropic and ahead of OpenAI on that front. All of those labs are way "ahead" of the cloud providers, but those providers are building infrastructure, not just training models, so it's not apples to apples.
1 reply →
> This is a pretty grim prognosis for European AI.
I think it's an incomplete read. What's the point in competing for a sizeable percentage of your funding when the finish line is incrementally being moved each month? Better spend it on leapfrogs which they seem to have done.
Meanwhile Mistral have a natural ace in their pocket with respect to regulation in the form of CADA and the Cloud Sovereignty Framework. I can't think of another company that would qualify as SOV-3 under that regime
So we’re supposed to cheer on companies that make worse products and only exist due to regulatory capture now?
Political polarization is turning the world insane.
I thought they wouldn't release anything at all, it's not dead yet :)
I don't think it will matter in a year or so. We are clearly topping out on useful intelligence for an increasing amount of tasks, as demonstrated by more and more models reaching the "useful" barrier.
This barrier is not going to start moving dramatically. It will simply be mostly satisfied for most work we do. Mistral is going to get there, soonish, long before the economy takes an entirely different shape (in so far that even happens).
There will be super human intelligence tasks, tasks truly constrained by intelligence for quite a while. Those will be few and far between, relatively speaking. Mistral will have plenty of opportunity to capture the other stuff, with a fraction of the resources required that it took the frontier labs to get there first.
People were extremely dismissive of chinese models until recently. They went from 1 year behind frontier to 6 months behind frontier to 3 months behind frontier extremely fast.
To be clear I'm not trying to dismiss European AI. I am a proponent of it.
But Chinese models have very much earned their place. The same cannot be said of Europe, so far.
Serious people haven't been very dimissive of Chinese models since at least DeepSeek-R1 in January 2025. Sadly, we in Europe are much behind.
Important to point out that these 'X months behind frontier' really refer to the public frontier, and not the actual frontier, which private companies are free to protect indefinitely. Perhaps open models are in actuality 18 months behind the actual frontier - how would any of us know?
People are cheering for a kid that is gaining ground in an ongoing race.
Curious, why do you say the model is mediocre? I haven't tried it, so I can't pass any judgement... I've learned to distrust benchmark rankings. Are benchmarks and Artificial Analysis the yardstick you're using?
> Curious, why do you say the model is mediocre? I haven't tried it, so I can't pass any judgement.
It's nowhere mediocre.
It's toes-to-toes with GLM-5.3 which is one of the best Open Weight model available (With Kimi K3) for general reasoning.
I just runned it on code reviews right now and it was able to catch some thread safety issue than DeepSeek-4.1 didn't. And DeepSeek-4.1 is by no means a bad model.
Doesn't matter, it's excellent, and it's European with European inference, which solves the pains of all my clients trying to build data lakes and processes on top of it.
Nobody in the real world cares about minor benchmark differences in money losing coding agents.
And nobody in the real world is giving Altman or Musk their data.
I wish you were true - but I still meet a lot of people handling sensitive data and using free or cheap version of ChatGPT or Claude with their customer data.
I think we will see some horror stories come out with data leak in the next years.
Neat. Wait 3 months for the landscape to change entirely.
LLM development is jumpy. It’s hard to extrapolate very far ahead.
I agree.
When Europe does surprise us, I will be the first to commend their progress. But until then, this is where we're at.
Reminds me of Gemini 3.5 Pro
First few models will always be slow improving and worse. The way to improvement is working your way through a gajillion evals [1], finding bugs, gaps, and curating training data (this part involves human design as well as raw inference compute) to fix it. This is very time intensive and can't easily be "done once and then everyone has lesser work to do" since every model is different. Well, one way to accelerate it is to simply have more compute, which mostly openai and anthropic have[2].
This is mistrals first 1T-scale model and I expect the 4th or 5th generation to be close to the best for many purposes.
[1] These evals differ from the public ones like terminal-bench, are sometimes model-specific, need real, diverse usage to actually create, and are held secretly since quality of eval is the first driver behind the next step improvement of a model.
[2] It is not close. This model was trained on less than 4k GPUs, whereas astra used north of 100k GPUs.
Mistral is not that new a player though. How can we give them this much grace when other players like xAI have done more in even less time? I don't think coddling Mistral helps them.
And to the point of scale and training cluster, so what? Not only do Chinese labs have smaller clusters with less empowered GPUs, compute is Mistral's responsibility. You can't take away from other labs just because they fulfill that responsibility better.
2 replies →
hurr durr
Those are comments from Europe. The US is waking up now and I expect them to be much harsher.
I really want them to win as that's our last horse in the AI race, but ~200 research-oriented devs out of 1800 employees? I believe they agree it's pretty doomed and have pivoted.
What? The US has been awake for almost 7 hours now
It's a bit below DeepSeek 4.1 Flash at about twice the size. For a model that was supposed to come out a few months ago this is pretty good. Mistral catching up to the chinese open-weight models is great news. Excited to see how they will build on that!
The blog post https://mistral.ai/news/mistral-large-4/
Very nice progress. Also I respect putting Kimi on those charts. Regardless of if they are beating the Pareto frontier (not now), model diversity is a good thing for humanity — I’m hopeful for the team to keep increasing their gains.
Open weight, European, competes with GLM-5.3 on cybersecurity. What's not to like?
[flagged]
I'm literally English and I feel the need to defend the French here, they're among the most successful military powers in the world historically speaking. It's hardly fair to judge a thousand years of French military history by the outcome of one war where they didn't do well.
1 reply →
Europe needs a lot of these. Quickly. Way to go, Mistral! Keep 'em coming.
I’d rather have one European AI lab like Mistral, with the financial firepower and compute to pretrain its own large models, than 20 weak labs that fine-tune Chinese models.
AI needs big bucks, and Mistral is Europe’s Anthropic.
just curious, how do we know that le chonk isn't just a fine-tuned chinese model? and/or distilled from US models?
4 replies →
Two is enough for foundation models, these guys and Aleph Alpha. The gap is closing.
You have to thank the US investors who funded Mistral from the very beginning.
Mistral would have gotten a tiny and measly "EU grant" and ASML would never have invested later had it not been for the US VCs.
Kudos to the US investors for being such selfless benevolent charities.
thank you for this and all the other great gifts the US and its companies have given to the world! we all love the us!
1 reply →
Ah yes. Let’s thank the might US for providing pesky Europeans capital a fistful of dollars. But maybe let’s do it after Americans thank for Russian, Arabian, Chinese, European capital and workforce. After all this is what “Made in the USA” means.
> Europe needs a lot of these.
Europe needs profitable AI companies, not money pits.
The US has made sure that Mistral has a large market in the EU by temporarily preventing non citizens from accessing Fable.
A lot of European companies now want a model the US can't cut off, but also lack trust in Chinese models.
Some of these will self host Mistral but most will pay them by the token. It's not going to be a huge market or a huge margin within that market but probably it'll be enough.
This is the exact mentality that makes the EU fall behind. If you don't want to invest in something until it makes a profit, you don't get the benefits of being a pioneer.
3 replies →
> Europe needs profitable AI companies, not money pits.
AI is a strategic technology with obvious national security implications. EU should invest in its development whether it's currently profitable or not.
1 reply →
there are no profitable AI companies at the moment ... This is the exact problem of European startups, trying to make them profitable from day 1 while American counterparts (and Chinese) keep bleeding money for years. Europe will never have a Tesla, a Google or an Amazon with that mindset.
4 replies →
Name 1 AI company that is profitable, and no Meta and Google are not AI companies
1 reply →
Like the US!
…oh…wait…
This is awesome, one of the coolest Pokemon ever too for those that don't follow that universe :)
https://bulbapedia.bulbagarden.net/wiki/Lechonk_(Pok%C3%A9mo...
Was wondering what Le Chonk meant (not French).
It's a joke, chonky is used to refer to fat cats, and there was a joke meme over the summer that Mistral are going to release a new model, Le Chaton Fat (chaton is kitten in French). The name is a nod to the memes.
3 replies →
The main thing I always get away from the comparison tables of these "big" models, is how well Deepseek v4.1 Flash performs. While still being the cheapest model by a long shot.
Beyond benchmarks, does it in day-to-day? Have always struggled to get competitive performance out of any Deepseek release going back to V3 vs Z.AI and Moonshot models. Maybe I really suck at whatever is needed to make DS models fly, but even tailoring my suite hasn’t gotten me far when I tried with V4 Pro. Happy for anyone who is able to leverage their models well, wish I’d be able to crack how to leverage them.
Will say their research is some of the best reads in the industry and I could not care less about their model release cadence as long as papers keep coming.
Is GPT Luna 6 dethroning Deepseek V4.1 Flash? It's price seem to be undercutting flash at a relatively similar capability.
-5 on omniscience? https://artificialanalysis.ai/evaluations/omniscience
That's not particularly great.
That said I love that they don't seem to restrict cyber capabilities to any degree and even lean into it.
If the best model for cyber attacks is open for everyone to use it just makes us all safer I think. Of course you then also HAVE to use it or otherwise you're vulnerable, which is a great distribution play.
I ran it on my trivia game Redactle which features a redacted Wiki article. Mistral Large 4 is not very good. It can sometimes solve a game with ~40 guesses whereas the top models like Gemini 3.8 Flash or Grok 4.7 can one shot most puzzles. My benchmark here aligns with AA Omniscience. I also have a version where the text is rewritten to detect over fitting to exact wiki text which changes the scores but not the leaderboard order.
https://redactle.net/llm-leaderboard?view=vital-500
Gemini models are summarizing wikipedia articles all day (when being used for Google's ai answer), can we draw some conclusion from this, did they train it extra well on wikipedia content?
https://artificialanalysis.ai/models/mistral-large-4 for the main stats
Imo omniscience correlates better to how useful the model is in practice than the intelligence index. But you have to use both together of course.
1 reply →
>If the best model for cyber attacks is open for everyone to use it just makes us all safer
Issue is..
I don't believe for an instant that any of us, including US citizens, get access to the best models for cyber that the US has. I think any adversary would have to assume the models in use by the US side are unreleased.
US is not the only one dealing under the table by the way, I also think everyone should take China having unreleased models as an operating assumption at this point.
So Mistral is the best that the public gets access to. And that's if it's even the best? Benchmarks and pragmatic work have often been shown to be two radically different things in this industry.
If the best model for cyber attacks is open for everyone to use it just makes us all safer I think.
The NRA approach to AI safety.
Seems hard to get customers at that price range when you're competing with open source models that are 1/2 - 1/3 the price but with similar capabilities.
People usually buy the cheapest, like Deepseek or GLM or they spend on Anthropic/OpenAI subs.. Are these models in the middle getting any users?
On a side note, I wonder if this was the popular free Space Bunny model that left openrouter yesterday.
If you are a EU company worried about your data then mistral is your only option. Think about EU military companies. They can't use US and Chinese models.
why couldn’t people just use chinese models rehosted in the EU? They’re open weights so anyone can serve them for any data residency requirement
5 replies →
I think you misunderstand how data processing works in a LLM. You absolutely can download the weights of a chinese model and run it on hardware you control.
2 replies →
Excited to hear this!
I barely use anything outside of cheap Chinese models on OpenRouter anymore. They are simply (more than) good enough for most of the things I do.
This model looks reasonably cheap. Though not deepseek levels.
Going to test it with Hermes, wondering where it will land in term of capability.
Bon chance, Mistral!
I'm at a loss as to what to do now. I've been wanting to support Mistral for so long. I struggled on with Mistral Medium 3.5 for longer than I should have (although I also think it taught me some valuable process lessons).
Recently I switched to the Mistral hosted GLM-5.3, this worked very well and powered through a tonne of work. Unfortunately, I also completely maxed out two subscriptions within the space of six days this month. One can't stack subscriptions with Mistral, so I'd have to register a third account for another subscription, which will be annoying with changing API keys all the time. Sure I can switch to pay-as-you-go API, but that adds up really fast. The Mistral dashboard shows that a Vibe CLI monthly subscription for €18.44 actually provides €255 worth of API use (apparently, and I tried to check this with Support but it seems like they were intentionally vague).
After maxing out my Mistral subs this morning, I dropped $10 on Xiaomi to try MiMo-2.6. So far so good, seem to have done a lot of work for the $2.85 I've spent, and Xiaomi prices are still much better than the Mistral introductory offer for Le Chonk.
Not sure where to jump.
Edit: Not being able to stack subs is my biggest gripe with Mistral. I'd probably pay them $100 per month (5 subs worth), but I'm not going to switch to the pay-as-you-go API and burn much more money for the same amount of tokens. Instead, I've taken that extra money elsewhere. If they just allowed one to keep topping up subscriptions on the same account it'd be grand. Or even a bigger single subscription. Make a $100 tier with five times the capacity.
You can confirm the allowed usage amounts with something like ccusage or tokscale
Yes, but I still don't know where the truth is.... The Mistral dashboard shows a drawn down on €255 worth of "credit" when you have a Pro subscription, which costs €18.44. There's no way I'm going to use the pay-as-you-go API if it means I'd be burning €500 in the next six days, as I just ostensibly have in the last 6.
I appreciate I can set a monthly spending limit for the pay-as-you-go API, but I'm just not willing to find out how far €100 will go, when I know it will go further elsewhere.
I just wish they had that €100 subscription tier, for the equivalent of €1k pay-as-you-go use.
Curious that it scores higher than opus 5.5 in cybersecurity because the closed models refuse to comply. I wonder if that means it's more susceptible to offensive uses.
The closed providers are serving useless, bricked models that are tainted with their shitty system prompts
One step closer to Le Chaton Fat.
Can't wait until Mistral starts adopting fancy product line names.
Soon we'll have Mistral 6 Chaton, Mistral 6 Guépard, Mistral 6 Tigre, Mistral 6 Dents-de-sabre, Mistral 6 Beast King, etc.
Panther, Jaguar … Wait, wrong company.
Le Chaton Fat development was cancelled internally after it broke loose into the treat box.
Refreshing to see this.
The pricing ($.68 in/$.07 cached/$2.09 out) makes it much cheaper than Kimi K3, GLM 5.3, and Meta Muse Spark 1.3. That's great!
But also much more expensive than GLM 5.3-flash and Spark 1.3 Contributor (the Meta-takes-your-data pricing of Spark 1.3).
So, I think it would have to be significantly better than GLM 5.3-flash to be worth it. GLM 5.3-flash is already very good.
You are quoting the discount pricing. It is 50% off for the next two weeks. The blog post has the real pricing up front in the card on the right side: https://mistral.ai/news/mistral-large-4/
After that, it will be much closer to GLM 5.3, but you can also get 5.3 in their API! I dont see people really talking about that.
Good point.
I still need to evaluate it for my own workloads, but if you trust the benchmarks, seems about on par in quality vs GLM 5.3 Flash
Source https://artificialanalysis.ai/models/mistral-large-4?total-c...
That is the promotional price, after which it doubles :(
I've been dreaming of this for a simple reason: the french prose combined with GLM 5.3 reasoning capabilities.
GLM 5.3 is incredible because for the first time with an open-source model, it feels.. enough. I don't need much anymore, this model is great in everything. Except a thing : speaking french.
If the benchmarks are true, I'd be glad to switch entirely to Mistral.
Excited to try this. The low costs v. benchmarks alone here are worth a serious test. K3 has been my daily driver for a month or two now and it's dramatically reduced token spend (while not having much of a negative impact on productivity).
This was the era of the AI race I was waiting for.
Curious now that SOL is cheaper than k3 - is k3 still your primary workhorse?
Haven't tried it yet. It looks like sol is a hair more expensive on input, a hair less expensive on output ($2/in, $10/out per 1m, K3 = $0.82/in, $13/out per 1m).
How could I resist switching to a model named after my cat!?
Stats be damned irrelevant. The naming is good with this one!
you and 10k other redditors
Le Chaton Fat is here!
Le Chonk https://www.youtube.com/watch?v=hD51W2txi1Y
Yeah just saw that, I'm gonna keep converting it in my head.
Maybe I'm missing something. Doesn't seem super impressive to me. A proprietary model with performance comparable to GPT-6 Luna and Deepseek 4.1 Flash, but at a higher price than either. The main selling point is that it's made in Europe... not very compelling, globally. I suppose maybe there is some niche where European-hosted open-weight models aren't enough to satisfy some EU regulation, where only the use of European-trained models is in compliance, but as a non-European I have no idea what that niche would be.
Side note: Wish this thread was more focused on talking about the model instead of debating about China and America. Whatever happened to staying on-topic?
You yourself also kind of pointed out why the discussion about USA and China is not off-topic. The niche Mistral wants to fill (afaik) is that it's European. I, as a European am really happy that they made such progress in so much worse (financial) conditions. I guess it's mostly geopolitics.
It's proprietary only until the end of the month, when its weights will be released.
Awesome!
All things considered I'm more inclained to pay for EU-based AI in the end (supporting local company and most likely being more aligned with EU regulations…)
Nice! Once they make it available through their API I will be happy to support them. My local Qwen3.8 27B is serving me well, but I miss the speed and concurrency that comes with subscriptions, and I am not currently paying for any.
Tais-toi et prends mon argent!
I believe it is already available in the API no?
Direct through them?
2 replies →
I don't know why, but I personally find Mistral's marketing strategy much more appealing than that of other companies.
For example, there's something about Anthropic's picked design and their little Claude avatars that's unsettling to me.
Sans context, I really like the Anthropic and Claude's faux-academic, minimal-ish, intellectual-ish branding.
(Especially the way it looked ~12 months ago -- it's gotten more cluttered since then. Perhaps unavoidably, as the breadth of their offerings has grown)
But over time it's begun to feel like unsettling cognitive dissonance as their ambitions grow and the stuff to worry about has piled up.
> faux-academic, minimal-ish, intellectual-ish branding
By this you sourely don't mean the messages shown in Claude Code, where Pi would show "Working..."
Academic style: Cooking... Sautéeing... Julienning... and similar annoying faux-brogrammer moody status messages.
2 replies →
The logo isnt a stylized butthole. That sure helps endear me to them.
I know the whole “AI logos look like buttholes” thing is a joke for most people, but I really speaks to the pornification of our society. I would never have thought “butthole” looking at any of their logos.
3 replies →
If it looks like that you, you will be seeing it everywhere.
2 replies →
Why was I about to write this exact same comment, and why did you beat me to it?
You are assuming that OP is not a fan of buttholes
The entire Anthropic branding is religious kitsch - deeply off putting, but apparently quite reflective of their reality.
I think there's such a thing as throwing too many marketing and sales people and too much polishing and "refinement" at something. I see in new product announcements from Microsoft as well. It's like seeing someone try too hard to impress you.
They've polished all personality out of their companies.
And their cookie banner. Never thought I would like a cookie banner
I do my best to avoid any AI marketing because I extremely despise it. I just haven't quite decided on my new profession, yet, but it's either going to be with plants or with animals.
Well, with this and Beam people are going to have to stop saying that western open models are dead. This is great news. Anthropic and OpenAI may have a bit more knowledge, talent and compute but they don't have a monopoly.
Investors looking for them to make monopoly profits are going to be disappointed The premium they can extract from consumers for their models will be capped. Tokens are likely to remain close to the cost of compute, a cost which is high but falling fast.
Also, yay Europe! Although the comparison between this and mimo 2.6 is not flattering...
Still second most expensive open-weight model. I don't care about cybersecurity index. And still can't beat Chinese models but good to see European in the game.
From Guillaume
> The RL run behind this preview is still in flight and shows no sign of saturation -- we will release a final version before the end of the month along with the weights of the model.
https://x.com/GuillaumeLample/status/2107461898127954001
https://xxcancel.com/GuillaumeLample/status/2107461898127954...
Mistral Large 4: 1050B, 49 Active
GLM-5.3: 753B, 40 Active
I was hoping for something that hinted at smaller models too, but I guess not.
Any competition is still good, especially now that the USA AI labs are starting to do regulatory capture.
Give them time. ML 4.0 was just pretrained. Mistral will certainly use it as the base for distillation and RL for smaller, better, more efficient iterations, just as the competition does.
It's competitive!
Good enough to show competence, and instill confidence in the team/company. Later releases can be more efficient.
I think it's a great release with that framing.
It's Mistral Large, they usually publish Medium and Small later
just keep RL frying it should get better...
Off Topic - The Mistral website - Really nice design. My guess, built by a human.
The key factor isn't whether the author used an LLM, but whether they had taste and attention to quality and iterated accordingly.
My guess would be designed by a skilled human with help from AI and built with AI by a skilled programmer.
Something looks off in artificial analysis. Benchmarks aren’t everything, but not even close to the Pareto https://artificialanalysis.ai/models/mistral-large-4?cost=in...
I guess lots of token usage.
Claims to be on par with GLM 5.3 in DeepSWE (from https://thenextweb.com/news/mistral-releases-large-4-a-1-tri...)
Title is missing the official model name (Le chonk)
We're probably fast approaching the scenario where the cheapest models will win out.
Is that a reference to LeChuck in Monkey Island? Love that game!
No, it's a play on the Twitter/Reddit memeing about "Le Chaton Fat", see: https://www.reddit.com/r/MistralAI/comments/1u6f0dm/what_is_...
You'd be surprised how much of the AI world is fueled by memes.
litterally, that's what an LLM is doing?
I think it's actually a wink and a nod towards the social media meme of "Le Chaton Fat," a fictional model that is jokingly attributed to Mistral, usually with century-defining benchmarks and unfathomable size.
https://x.com/i/trending/2066326562623127678
Why is distillation weird?
Benchmarks are better than expected! And probably got there without distillation ;)
Is there a reason to believe why they wouldn't distill locally running open weights Chinese models?
Looks like they are doing 50% off to stay price competitive with DS Flash V4.1
I tried Mistral's last 3 large models, devstral and mistrallarge3 and the numbers were not even benchmaxxed, but just false. the models were so weak and garbage. Let's hope they are telling the truth this time around, we need more alternatives. There mistral-small and original MistralLarge models were awesome, hoping they are back!
Excited to test this on my benchmark[1], but I'm guessing it won't fare better than MiMo v2.6 Pro which is currently the best open-weight model as far as I'm concerned.
Why isn't this on openrouter yet? Is there a better router that gets models much quicker?
1 - https://bench.killswitch-lang.org/
https://openrouter.ai/mistralai/mistral-large-4-0
Good to see Europe is at least a little bit still in the game.
https://docs.mistral.ai/inference/model-selection-guide?mode...
Cost is stated at half the price of GLM-5.3, which is quite interesting.
The terminalbench 4.0 score is encouraging as a sign of it not doing anything "stupid" when put in a proper harness.
I tried a quick abbreviated kimbench on their playground before bothering to do the whole thing.
Maybe I didn't really select mistral 4? Either way, failed completely on question 1 and the next 2 questions were completely off base too. I didn't bother to finish.
Not suitable for my purposes I don't think.
At the end of the blog post we get this nugget.
> The pace of progress from here will be fast. Stay tuned.
Is the reason for the massive gains in certain benchmarks due to distillation from the other lead models hence the slightly "under" pattern seen in the comparison charts?
no hugging face link :( ... but hey its on openrouter yay
Here are my results
https://dach.peerbench.ai/compare?models=mistralai%2Fmistral...
Looks like a bit better than the recent Kolibri-1 but still below Qwen3.8 27B
https://venturebeat.com/technology/mistral-debuts-large-4-le...
I tried testing it, but reasoning effort parameter doesn't seem to be there and output is sort of broken because it reasons directly in the output tokens...
Awesome!
All things considered I'm more inclained to pay for EU-based AI in the end (supporting local company and most likely being more aligned with EU regulations…)
Excited to see this! Nice that they are saying this is just a first step.
Give them more compute!
What experience have people had with Mistral models for cyber research? are they as constrainted as Anthropic models? I use claude day2day but have the need for a model with fewer guardrails to pentest our own APIs.
They explicitly lean in to cyber work, and they appear to be very permissive from their marketing:
> This is particularly important in cybersecurity, where provider-level refusals can block legitimate vulnerability research and incident response, and where losing access to a capability mid-incident can itself become a critical security risk. ML4 pairs top-tier cyber performance with open weights and self-deployment, giving organizations both the capability and the autonomy to run advanced security work under their own policies.
> That top score reflects a practical advantage. Several leading closed models, including Claude Opus 5.5 and GPT-6 Astra, score near zero on the same test because they refuse to perform the task. Yet defending software often starts with proving that a flaw is real, exactly the kind of work safety filters in closed models can block. This matters even more as threat actors increasingly jailbreak those same models to support offensive cyber activity
le chaton fat is real, my life is complete. Benches look crazy good for 1T.
Seems like they've finally made a genuinely competitive model since the original LLM craze, congrats to them! Glad to see some diversification in open weights providers.
What's up with the name? It reminds me of my teenage self trying to speak in funny memes.
Does it have the ability to capture the market like OpenAi or Anthropic ? My point is, Regular/Average users of AI do not really care about benchmarking. Marketing decides which company makes it to the phones or PCs.
> Marketing decides which company makes it to the phones or PCs.
Right now a big part of LLM market is people using it for professional software development. Most of these users probably care about the quality of the model and also notice it during daily work.
For normal consumers, shure it doesn't matter. In the end the ai summary of google will probably be the most used as they are already exposed to it anyways.
dont take it personally, i just dont understand why to release a model that is not showing new strong capabilities, why would anybody use this model and not Claude Opus.
Because screw Anthropic and OpenAI, that's why.
I love the name! Teasing the ones making fun of them.
Without exaggeration, given a choice between models, I would pay for Mistral's model over Anthropic's based on the name alone, completely ignoring features or other technical considerations. The name is playful and is such a refreshing contrast to Anthropic's (and OpenAI's) doomsaying, scaremongering, and god-posturing.
This sure is nice, but I've had less than satisfactory results with GLM5.3. I'd like Mistral to compete with Qwen3.8-Flash-Next a 120B class model that IMO is the first model that I can use for serious coding while running it locally.
I estimate it's coding ability on par with opus 4.6 (but opus definitely beats it on factual knowledge) Still it's a genuinely useful model, when everything else except Anthropic's models (and for only 3 weeks after it came out Google's Gemini 3 pro, before it got merged) are not to me.
I'd live to have one like that but EU made.
I really hate this open weight but closed science approach. These companies just take from academics and all the Chinese companies that are doing good science, but without understanding the recipe it makes it hard to know where the failure points will be until your agent accidentally commits a crime.
Mistral doesn't publish the science.
Brother I think No one publishes the science
if it's not available yet why have a 'try it today' header at all?
> "Try it today" > > There is still more to come. As we work toward releasing the weights, we will share further details on the model architecture, additional benchmarks, and our post-training methodology.
You can try it on their website, https://console.mistral.ai/playground
It's just that the open-weights aren't yet available (although the long delay is slightly annoying).
Since the Chinese companies publish their research it would have been odd if Mistral didn't start catching up.
It certainly has helped OpenAI and Anthropic get their KV cache costs under control.
It's no secret that everyone is dis-stealing from everyone else.
I don't see how distillation relates to using the published techniques developed by Deepseek, Moonshot, Zhipu, etc
touche
> Trained from scratch
How are they training without pirating the Z library corpus and all that?
I interpret "scratch" to mean brand new weights. Not that they aren't training on a corpus of human text
Right but how did they get a corpus, how do they compete legally without distillation?
1 reply →
at 200M tokens for the full AI suite run its not token efficient at all
Glad to see progress, despite the ever-increasing sabotage by the EU bureaucrats
So about 2 or 3 generations behind, just like they were a year ago?
Half price on open router right now
sorting the charts like that gives off weird vibes
https://mistral.ai/news/mistral-large-4/
sorting the chart like what? You just linked to the main page.
If you scroll down, most of the charts on that page are sorted s.t. Mistral's bar is right next to the worst competitor model, while the best competitor model's bar is positioned on the opposite side.
If one were to be cynical one could say that it's intentionally making Mistral's result look better than it actually is by making it harder to compare the bar heights.
Number one in Sovereign AI. Join our Discord.
We actually got Le Chaton Fat before GTA 6
Is there consensus on if this was https://openrouter.ai/stealth/space-bunny-alpha ?
Space Bunny Alpha is probably MiniMax M3.1 (rumors on Twitter since it seems to have a similar tokenizer).
Yes broad consensus had developed in the ten minutes between announcement and you asking if consensus had developed, and I'm excited to report that it consensed in the affirmative -- it IS Space Bunny Alpha!
sorting the charts like that gives off weird vibes
Had the same thought - feels chart crime adjacent
Bar charts should start at zero. If they don't start at zero, there should be a clear visual indicator that the chart has been trimmed without having to read the axis labels. I hate that this has to be repeated so often that it has become a cliché.
I thought lechonk motto was just a meme!
I'm glad they're keeping at it!
Previous one is barely in top 50 on arena.ai
A bit disappointing to see it still lagging behind Chinese open models. Those Chinese models are pushing proprietary models to raise the bar, but we need equally strong non-Chinese open models to challenge the Chinese ones in turn.
I mean no disrespect but these are terrible numbers or am I missing something? It seems like Mistral continues to only be relevant for people that want a model trained in Europe. Too bad.
On Prem. Thats a bid deal for some enterprises.
also the benchmarks are not necessarily indicative of how well the model will perform in its own harness with its own skills.
Any open weights model is "on prem".
I mean usually the benchmarks make any model feel better than how they actually perform.
better than K3 and DS4, cool
Amazing! just tested
impressive release this time by Mistral. bullish.
Not bad only two major releases behind top tier. Edit : checked its rather 3 generations behind . Oh well
Looks like it's about a year behind still. i.e. its intelligence is behind models from roughly a year ago.
https://www.vals.ai/benchmarks/vals_index
Where can it be tested?
where does sit on the pareto distribution compered to Le Chaton Fat?
1T parameters -- ugh, open models keep getting bigger and bigger! Running them at home is getting ever more unattainable, especially for those of us with bandwidth-poor hardware like Apple silicon -- please continue releasing smaller models, too!
Massive fumble not to call it “le chaton fat”.
>> Unofficially ML4, very officially: le Chonk
Honestly just nice to see a leader in this space not take themselves so seriously.
Too bad this got marked as a dupe, as it actually has benchmark info unlike the other page which is just docs.
The weird thing is how worse they are at things like coding than other open weights. You'd expect them to at least distill coding from other open weights to match them.
Pretty impressive. I genuinely wonder how Mistral hires talent when their salaries are so terrible. Guess there aren't many better places to work in Europe.
europe finally getting into the race here.
Awh I was half expecting a zombie pirate..
wowza le models a heckin chonker
YES finally
Anyone have any indication when I can get my hands on a developer plan for this?
They do sell subscriptions I think
Don't believe Mistral. They're wrong. It's really called "Le chaton fat".
Also 1T-A49B. Weights currently closed but promise to open source them by the end of the month.
Great release movie.
>Don't believe Mistral. They're wrong. It's really called "Le chaton fat".
OpenAI's therapist: Le Chaton Fat isn't real and cannot hurt you
Le Chaton Fat:
Woah, this seems like a big deal (assuming the benchmarks are as good as claimed)?
Mistral slightly proving me wrong (and I'm not mad).
Now THAT'S how you name a model. Take note, others.
>Frontier performance
* proceeds to not compare to Opus 5.5
They say "open weights frontier performance". Of course they aren't comparing to Anthropic, because they are not putting their weights online.
I'm talking about the section of the article labeled "Frontier Performance", which does not specify that.
quick question why put GLM 5.3 at 61 while a quick check on DeepSWE 1.1 puts it at 69?
also they forgot muse spark at 75% while claiming they were outshining all US models?
Can we consolidate the posts? Currently there's 3 on the front page, basically all pointing to Mistral's messaging in different places.
[dead]
[dead]
[flagged]
[flagged]
[flagged]
[dead]
[dead]
[flagged]
[flagged]
To Europeans at least, this is a big deal, especially because of cybersecurity capabilities.
It is of vital strateigic importance for Europe (and really the rest of the world too) that there are non-American, non-Chinese options for AI.
Non-American maybe, but non-Chinese is likely impractical. We might not want their APIs, but we can't compete on energy and labour for training, even if it means buying licenses to host the weights (which CADA encourages). From China's perspective, what's the case for baking weights they can't sell? Some Chinese models already don't even have the Taiwan politics stuff baked in, those filters are only in their APIs.
I realised after writing this, the EU as is uncomfortably often the case, may be the real forcing function for what happens with US policy irrespective of the media campaigns we're presently seeing. Here's hoping for a steady trickle of stale ChatGPT weights leaking from EU infra providers in the long term.
5 replies →
What makes you think the rest of the world trusts Europe more than the Chinese?
1 reply →
Mistral was silent on the frontier for quite some time so this is exciting to me.
It's hacker "news" not hacker "scientific breakthrough"
Mistral is the only major non-Chinese contender in the AI race releasing open-weight models. I’d rather trust that French weights haven’t been backdoored than Chinese ones. Heck, I’d even trust Mistral more than Anthropic for sensitive work.
It's the only non-Chinese open weights frontier model, so while not groundbreaking per se, still quite important.
Important for sovereignty, multi-language support, and choice.
You have a number of other European providers of open-weight LLMs, such as Bielik and Aleph Alpha. They are not trying to compete at the frontier, but they sometimes develop original architectures. I am also a big fan of PleIAs' research in this regard, have a look at their Baguettotron
1 reply →
There are many non-Chinese open weights models, just not very good ones.
2 replies →
Inkling!
[flagged]
Mistral wont win the AI race because of the model names. I wont bother an arrogant Parisian hipster with my insecure prompts who then plays with his moustache and responds with a judgmental "pfff"
i subbmitted a partnership proposal in your contact.
Does Mistral ever advance the state of the art on any dimension?
And if not, why do they exist?
Update: The number of people advocating not innovating is wild. There is no reason why Mistral cannot innovate in ML, they explicitly choose not to. My point is that, given that choice, they should spend their GPU hours differently.
"Sovereign AI" is a joke, there is no substantive difference between a post-trained open weight model from an American or Chinese company and what Mistral is doing today, beyond spending 80% of their GPU hours reproducing a last-gen model's pretraining.
Why would the Europeans want a high-quality open source model that isn't either owned by American Trillion-dollar companies or created by the Chinese? You really can't think of any reasons?
> high-quality open source model
> that isn't owned
I don't know what definition of open weights you are using, but the "open" part is what allows you to use it without being beholden to American Trillion-dollar companies or the Chinese.
Mistral could do exactly what Cursor does, and post-train the model to meet the needs of Europeans.
That's a far more efficient (and useful!) use of GPUs than yet another pretrain on an outdated model backbone.
Has the French military advanced the state of the art on any dimension in the last hundred years?
And if not, why do they exist?
Second hand rifle distributor?
Of course they have!
1 reply →
So that there exists an EU-native option in the near-frontier LLM space?
Not everyone is wild about being downstream of either the Chinese or US governments, particularly when it comes to things like cybersecurity
Just like Microsoft and most companies around PCs, their existence isn't about pushing any boundaries, but seemingly targeting deploying what's already been "invented" into the corporate world and bigger companies. Not "wrong", just different.
Sure, agreed. But they do their own pre-training, at great expense, on outdated model backbones.
Wouldn't it be better to do something more like Cursor, and RL on an existing pretrained model if you're not innovating anyway?
1 reply →
Mistral is one of the few European AI labs. Look up "sovereign AI".
Does erichocean ever advance the state of the art on any dimension?
And if not, why do they exist?
Yes, actually. Thanks for asking.
Finally, a model small enough to self-host on my 2012 MacBook Air if I don't mind my house reaching room temperature in 2 seconds.
You might be missing three 0s on the parameter count, or what am I missing?
I'm assuming you need somewhere 0.5 to 1TB of RAM for the weights only
not to ignore you but how is it possible that I have zero karma
1 reply →
Still a long way behind closed models sadly :-( The gap between open and closed seems the biggest it's been for a year or two.
I try to stay up on AI models, but I've given up on Mistral. I have tried using it too many times and it's the worst out of major models. Even Kimi and DeepSeek are miles ahead.
Seems like it's another European company that is only alive through government financing.
I understand wanting a home grown industry, but with AI models, just abliterate a DeepSeek model.
Total misunderstanding of the industry. Europe needs hardware, not a model that needs to be replaced monthly.
> Total misunderstanding of the industry. Europe needs hardware
With all due respect, I have no idea what you are trying to say with this comment and have no clue how "hardware" would turn Europe's tides. You explained nothing.
Do they need training hardware? Inference hardware? ASICs or GPGPUs? Edge hardware? Agent hosts? Faster cores, or wider ones? Taller systems, or more parallel ones?
Your vagueness completely undermines the authority that your criticism relies on.