Comment by mhaberl
4 hours ago
2 things here, one is related to benchmaxxing, another to being good enough
1. Mistral isn't benchmaxxing. That doesn't mean they're better, but it does mean the benchmark gap not a good reflection of the actual gap
2. I think the "world class or nothing" framing mixes general capability with system capability. Most deployments don't need AGI In RL you need a model that's reliably good at one or two things, thats it. Example: Case of a hospital flooded in emails. You make a system that decides which patient emails needs a human and drafts replies for the rest. If a sovereign model is good enough at that, and you can run it on a hospital's own servers under EU jurisdiction, the frontier gap part has zero importance
Who cares about "beats DeepSeek / GPT11 / Claude Fairytale 8.9"
Agree benchmarks don’t tell the whole story. A better and easier evaluation of capability is whether anyone is using it for anything.
Does Mistral have material market share for any application, including anything that would fall under item 2 above?