Comment by blfr
9 hours ago
I don't care about benchmarks. Benchmarks show that Opus 5 is a stronger model than Fable 5 which is obviously not the case.
But I do care about capability and so far only Anthropic and, very recently with Astra, OpenAI can deliver on coding quality. And capability matters immensely. There is a world of difference between being able to do something and not being able.
People have been using LLMs for two year. It's not just this week's LLM release that is capable something.
A capability isn't binary. There is a massive difference between can produce an impressive demo and can reliably complete the task without a human babysitting it.
Yup, new SOTA models especially with high/xhigh/max reasoning too often overengineer solutions, good for benchmarks that usually measure task completion, bad for normal development where you don't want 'rewrite in rust and 1k LOC unit tests' style solutions when agent does mundane bug fixes.
1 reply →
It's not this week's change. Fable was the step change for programming. And most of truly useful and powerful capabilities arrived in the last eight months.
> Fable was the step change for programming.
AFAIK: Mistral does not even try to compete in this field. There are other use cases for LLMs beside coding. As Mistral AI wrote:
> During the first wave of generative AI, the central question was who could build the most powerful model. Organizations and governments are now asking a different one: how to harness the power of AI for their mission-critical needs without surrendering control over the infrastructure and intelligence loop. Demand for that combination of performance with control, choice and independence is growing internationally, as enterprises and governments weigh the long-term technology dependencies, data governance requirements and deployment choices that come with any AI investment.
> Mistral is the only AI company in the world building the full stack required to answer that question: open-weight models, the infrastructure and the compute capacity they run on, and the products that bring them into production; ensuring that customers are never locked into a single vendor's roadmap, pricing or availability.
> Mistral’s full-stack and open approach also allows organizations to build on it without exposing their most valuable data, workflows and institutional knowledge to anyone outside their own walls. That's what makes Mistral’s stack the sovereign AI layer, meaning retaining control across four dimensions: data that stays inside the organization's boundaries, models that are controllable and customizable, compute that is private and predictable, and systems in production that are fully controllable and auditable.
1 reply →
People said this for Opus 4.6 too. Every release the models get RLHF'ed into accomplishing a new task and the people who need to do this task think there was a step change.
3 replies →
If the choice is between Mistral and no AI, I'll take Mistral any day
Even those old llama models were ok for coding
Yes yes they won't be like Claude's fire and forget (until you see how many tokens you burned to write "Hello World")
> don't care about benchmarks
You must care about good benchmarks (identify those that have relevance).
Genuinely interested, which ones do you think have relevance?
If I read forums and talk to people IRL most have differing opinions what model is best. Yes, for me it's pretty clear Opus is better than earlier models, but it's at least not obvious to me that the later are significant improvements.